PageSourceSearch

https://www.aievals.co/assets/governance--openai-preparedness-C2_8ERUp.js

js aievals.co collected 2026-10-08 06:46:02 UTC 21,154 bytes, 125 lines download raw bytes

1const e=`<article class="markdown-body"><p>OpenAI's Preparedness Framework is the company's public commitment to track and prepare for frontier capabilities that create new risks of severe harm. As of October 2026 the published text is Version 2, last updated April 15, 2025 <sup><a href="#user-content-fn-openai-preparedness-v2" id="user-content-fnref-openai-preparedness-v2" data-footnote-ref="" aria-describedby="footnote-label">1</a></sup>. It has the same broad shape as Anthropic's RSP (capability-triggered safeguards before deployment) with a different vocabulary and category structure <sup><a href="#user-content-fn-openai-preparedness" id="user-content-fnref-openai-preparedness" data-footnote-ref="" aria-describedby="footnote-label">2</a></sup>. This page summarizes the operational pieces and notes the differences a practitioner should know.</p>
2<h2 id="the-category-structure">The category structure</h2>
3<p>Version 2 tracks three categories closely and researches five more <sup><a href="#user-content-fn-openai-preparedness-v2" id="user-content-fnref-openai-preparedness-v2-2" data-footnote-ref="" aria-describedby="footnote-label">1</a></sup>:</p>
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25<table><thead><tr><th>Tracked Category</th><th>What it covers</th></tr></thead><tbody><tr><td>Biological and Chemical</td><td>Lowering the barriers to creating and using biological or chemical weapons</td></tr><tr><td>Cybersecurity</td><td>Scaled cyberattacks and vulnerability exploitation</td></tr><tr><td>AI Self-improvement</td><td>Rapid, hard-to-track acceleration of AI capabilities that could challenge human control</td></tr></tbody></table>
26<p>The Research Categories, where threat models or evaluations are not yet mature enough for tracking, are Long-range Autonomy, <a href="/glossary#sandbagging" class="glossary-link" data-glossary-term="sandbagging">Sandbagging</a>, Autonomous Replication and Adaptation, Undermining Safeguards, and Nuclear and Radiological. Persuasion, a tracked category in the 2023 version, is out; OpenAI now handles it through its Model Spec and usage policies. To get onto the tracked list a risk has to be plausible, measurable, severe, net new, and instantaneous or irremediable, and "severe" means the death or grave injury of thousands of people or hundreds of billions of dollars of damage.</p>
27<p>Each Tracked Category has two thresholds, High and Critical. Version 2 removed "low" and "medium" because they were not operationally involved in the work. A model at High must have safeguards that sufficiently minimize the associated risk before it is deployed. A model at Critical needs those safeguards during development as well, whatever the deployment plans.</p>
28<p>The thresholds are no longer hypothetical. In September 2026 OpenAI designated Astra the first model to meet the Critical threshold, in cybersecurity: in expert-led testing it found previously unknown vulnerabilities in a hardened browser and operating system and chained them into working exploits. OpenAI says it delayed parts of Astra's development and release and paused certain frontier training for two weeks to harden its research infrastructure before concluding the safeguards were sufficient for release <sup><a href="#user-content-fn-openai-path-to-astra" id="user-content-fnref-openai-path-to-astra" data-footnote-ref="" aria-describedby="footnote-label">3</a></sup>. Version 2 also says OpenAI expected to update the framework before any model reached Critical. The published text is still Version 2; in August 2026 OpenAI said it will evolve the framework to cover safeguards across training and deployment <sup><a href="#user-content-fn-openai-pacing-cyber" id="user-content-fnref-openai-pacing-cyber" data-footnote-ref="" aria-describedby="footnote-label">4</a></sup>. Check the live page before quoting thresholds.</p>
29<h2 id="how-it-differs-from-the-anthropic-rsp">How it differs from the Anthropic RSP</h2>
30<p>The two documents are structurally similar; the differences are in emphasis, and since Anthropic's v3 rewrite, in what each company promises to do on its own.</p>
31<p><strong>Categorization.</strong> The Preparedness Framework names Tracked Categories and evaluates each independently. RSP v3 names four 
31capability thresholds (two chemical/biological, one for misalignment in high-stakes settings, one for automated R&#x26;D) and, for each, sets Anthropic's own plan next to a more demanding industry-wide recommendation. In practice both end up with category-by-category capability evaluations.</p>
32<p><strong>Threshold language.</strong> OpenAI uses High and Critical per category. Anthropic no longer uses a ladder of AI Safety Levels for future capability; "ASL-3" now names the protections it has in force. The language matters for external communication: "High in Biological and Chemical" under one framework does not map cleanly onto "ASL-3 protections" under the other.</p>
33<p><strong>Governance.</strong> OpenAI's Safety Advisory Group (SAG), a cross-functional group of company leaders, reviews Capabilities Reports and Safeguards Reports and recommends; OpenAI Leadership decides; the Board's Safety and Security Committee oversees <sup><a href="#user-content-fn-openai-preparedness-v2" id="user-content-fnref-openai-preparedness-v2-3" data-footnote-ref="" aria-describedby="footnote-label">1</a></sup>. Anthropic has a Responsible Scaling Officer who approves development and deployment decisions, with policy changes approved by its Board in consultation with its Long-Term Benefit Trust <sup><a href="#user-content-fn-anthropic-rsp" id="user-content-fnref-anthropic-rsp" data-footnote-ref="" aria-describedby="footnote-label">5</a></sup>.</p>
34<p><strong>What a competitor's behavior changes.</strong> Both frameworks let the other labs move the bar, in opposite-sounding ways. OpenAI's marginal-risk clause allows it to relax safeguards in a category if it can rigorously confirm that a competitor developed or released a High or Critical system without comparable safeguards, provided overall risk does not rise meaningfully, the change is public, and OpenAI stays more protective than that competitor. Anthropic's v3 commits to delaying development and deployment in two cases: when it is clearly in the lead, and when all comparably capable competitors can show their risk is contained, in which case the delay lasts until Anthropic matches their safety posture. It treats a pause outside those cases as discretionary. Neither is an unconditional pause.</p>
35<p>For a practitioner copying the structure into an internal policy, both work. The choice of vocabulary depends on which language your customers and regulators already use.</p>
36<figure class="not-prose" style="margin: 2rem 0;">
37  <svg viewBox="0 0 720 320" role="img" aria-label="Comparison of OpenAI Preparedness Framework version 2, which scores three Tracked Categories independently against High and Critical thresholds, with Anthropic RSP version 3, which sets four capability thresholds and pairs Anthropic&#x27;s own plan with an industry-wide recommendation for each." style="width:100%;height:auto;display:block;" xmlns="http://www.w3.org/2000/svg">
38    <title>Comparison of OpenAI Preparedness Framework version 2, which scores three Tracked Categories independently against High and Critical thresholds, with Anthropic RSP version 3, which sets four capability thresholds and pairs Anthropic's own plan with an industry-wide recommendation for each.</title>
39    <g font-family="&#x27;Instrument Sans Variable&#x27;,&#x27;Instrument Sans&#x27;,ui-sans-serif,sans-serif">
40      <rect x="24" y="20" width="330" height="256" rx="8" fill="hsl(var(--bg-elev))" stroke="hsl(var(--rule-strong))" stroke-width="1"></rect>
41      <text x="40" y="46" font-size="14" fill="hsl(var(--ink))">OpenAI Preparedness (v2)</text>
42      <text x="40" y="64" font-size="11" fill="hsl(var(--ink-soft))">two thresholds per Tracked Category</text>
43      <text x="40" y="106" font-size="12" fill="hsl(var(--ink))">Biological and Chemical</text>
44      <line x1="240" y1="102" x2="320" y2="102" stroke="hsl(var(--rule-strong))" stroke-width="1"></line>
45      <circle cx="240" cy="102" r="4" fill="hsl(var(--ink-mute))"></circle>
46      <circle cx="320" cy="102" r="4" fill="hsl(var(--ink-mute))"></circle>
47      <text x="40" y="142" font-size="12" fill="hsl(var(--ink))">Cybersecurity</text>
48      <line x1="240" y1="138" x2="320" y2="138" stroke="hsl(var(--rule-strong))" stroke-width="1"></line>
49      <circle cx="240" cy="138" r="4" fill="hsl(var(--ink-mute))"></circle>
50      <circle cx="320" cy="138" r="5" fill="hsl(var(--accent))"></circle>
51      <text x="40" y="178" font-size="12" fill="hsl(var(--ink))">AI Self-improvement</text>
52      <line x1="240" y1="174" x2="320" y2="174" stroke="hsl(var(--rule-strong))" stroke-width="1"></line>
53      <circle cx="240" cy="174" r="4" fill="hsl(var(--ink-mute))"></circle>
54      <circle cx="320" cy="174" r="4" fill="hsl(var(--ink-mute))"></circle>
55      <text x="240" y="204" text-anchor="middle" font-size="11" fill="hsl(var(--ink-mute))">High</text>
56      <text x="320" y="204" text-anchor="middle" font-size="11" fill="hsl(var(--ink-mute))">Critical</text>
57      <text x="40" y="228" font-size="11" fill="hsl(var(--ink-soft))">High: safeguards before deployment</text>
58      <text x="40" y="244" font-size="11" fill="hsl(var(--ink-soft))">
58Critical: safeguards during development too</text>
59      <text x="40" y="260" font-size="11" fill="hsl(var(--accent))">Accent dot: first Critical (cyber, Sept 2026)</text>
60      <rect x="378" y="20" width="318" height="256" rx="8" fill="hsl(var(--bg-elev))" stroke="hsl(var(--rule-strong))" stroke-width="1"></rect>
61      <text x="394" y="46" font-size="14" fill="hsl(var(--ink))">Anthropic RSP (v3)</text>
62      <text x="394" y="64" font-size="11" fill="hsl(var(--ink-soft))">four capability thresholds</text>
63      <rect x="394" y="84" width="170" height="32" rx="8" fill="hsl(var(--bg-elev))" stroke="hsl(var(--rule-strong))" stroke-width="1"></rect>
64      <text x="479" y="104" text-anchor="middle" font-size="12" fill="hsl(var(--ink))">Non-novel chem/bio</text>
65      <rect x="394" y="124" width="170" height="32" rx="8" fill="hsl(var(--bg-elev))" stroke="hsl(var(--rule-strong))" stroke-width="1"></rect>
66      <text x="479" y="144" text-anchor="middle" font-size="12" fill="hsl(var(--ink))">Novel chem/bio</text>
67      <rect x="394" y="164" width="170" height="32" rx="8" fill="hsl(var(--bg-elev))" stroke="hsl(var(--rule-strong))" stroke-width="1"></rect>
68      <text x="479" y="184" text-anchor="middle" font-size="12" fill="hsl(var(--ink))">High-stakes misalignment</text>
69      <rect x="394" y="204" width="170" height="32" rx="8" fill="hsl(var(--bg-elev))" stroke="hsl(var(--rule-strong))" stroke-width="1"></rect>
70      <text x="479" y="224" text-anchor="middle" font-size="12" fill="hsl(var(--ink))">Automated R&#x26;D</text>
71      <text x="580" y="132" font-size="11" fill="hsl(var(--ink-soft))">each threshold has</text>
72      <text x="580" y="148" font-size="11" fill="hsl(var(--ink-soft))">a company plan and</text>
73      <text x="580" y="164" font-size="11" fill="hsl(var(--ink-soft))">an industry-wide</text>
74      <text x="580" y="180" font-size="11" fill="hsl(var(--ink-soft))">recommendation</text>
75      <text x="394" y="260" font-size="11" fill="hsl(var(--ink-soft))">ASL-2 and ASL-3 now name protections in force</text>
76      <text x="360" y="304" text-anchor="middle" font-size="12" fill="hsl(var(--ink-soft))">Both gate on capability; neither promises an unconditional pause</text>
77    </g>
78  </svg>
79  <figcaption style="margin-top:0.6rem;font-size:0.8125rem;line-height:1.45;color:hsl(var(--ink-mute));">
80    <strong>Figure:</strong> OpenAI Preparedness Framework v2 vs Anthropic RSP v3 threshold structure: Preparedness scores three Tracked Categories (Biological and Chemical, Cybersecurity, AI Self-improvement) independently against High and Critical thresholds, with Cybersecurity the first category where a model (Astra, September 2026) was designated Critical. RSP v3 sets four capability thresholds and pairs Anthropic's own plan with a more demanding industry-wide recommendation for each; ASL labels now describe the protections in force rather than future levels.
81  </figcaption>
82</figure>
83<h2 id="the-open-source-eval-repo">The open-source eval repo</h2>
84<p>OpenAI publishes a public GitHub repository of frontier capability evals. It was renamed from <code>openai/preparedness</code> to <code>openai/frontier-evals</code> (the old URL redirects) and, as of October 2026, holds PaperBench (replicating research papers), SWE-Lancer (paid freelance software tasks with end-to-end tests), and EVMbench (smart-contract security tasks, added February 2026) <sup><a href="#user-content-fn-openai-preparedness-repo" id="user-content-fnref-openai-preparedness-repo" data-footnote-ref="" aria-describedby="footnote-label">6</a></sup>. MLE-bench, the machine-learning engineering benchmark, lives in its own <code>openai/mle-bench</code> repo. These are useful in their own right as eval harnesses; they are also a reference for the kind of capability evaluation a Preparedness-style framework produces.</p>
85<p>The repo is one of the better resources for a team standing up its own capability evaluations because the eval definitions are concrete: specific tasks, specific scoring, specific reproducibility metadata.</p>
86<h2 id="what-to-lift">What to lift</h2>
87<p>Four patterns are worth copying into an internal version regardless of which framework you <a href="/glossary#anchor" class="glossary-link" data-glossary-term="anchor">anchor</a> to.</p>
88<p><strong>Per-category thresholds.</strong> The OpenAI category list reads like a risk taxonomy, and its five admission criteria (plausible, measurable, severe, net new, instantaneous or irremediable) are a reusable test for what deserves a threshold at all. Even if your product is not at frontier scale, naming the specific risk categories you track and writing thresholds for each is more useful than a single global risk level.</p>
89<p><strong>Pre-deployment gating.</strong> Both frameworks make the safeguard a precondition for shipping. This is the load-bearing piece. An internal version that says "we will track these risks" without a gate is documentation; the version with a gate is a policy.</p>
90<p><strong>
90A written case before high-risk work starts, with a dissent attached.</strong> OpenAI's September 2026 guidance on safety cases for frontier training runs asks for a structured argument before the run, a written dissent from someone on another team, sign-off from several senior leaders who each hold a veto, and controls that fail closed, so a run cannot start with monitoring switched off <sup><a href="#user-content-fn-openai-safety-cases-training" id="user-content-fnref-openai-safety-cases-training" data-footnote-ref="" aria-describedby="footnote-label">7</a></sup>. The same four pieces fit a high-risk launch at any scale, and the dissent costs an afternoon.</p>
91<p><strong>Public commitments where possible.</strong> Publishing the framework is the strongest version of the commitment. Most enterprise teams will not, but publishing the existence of the policy (a one-page summary on your trust portal) is the next-best step and is increasingly expected in procurement conversations.</p>
92<h2 id="what-to-do-this-quarter">What to do this quarter</h2>
93<ol>
94<li>Read the Preparedness Framework and the RSP back to back. The structural similarities make the choice of which to anchor to a vocabulary decision, not a substance one <sup><a href="#user-content-fn-openai-preparedness" id="user-content-fnref-openai-preparedness-2" data-footnote-ref="" aria-describedby="footnote-label">2</a></sup> <sup><a href="#user-content-fn-anthropic-rsp" id="user-content-fnref-anthropic-rsp-2" data-footnote-ref="" aria-describedby="footnote-label">5</a></sup>.</li>
95<li>Pick the three risk categories most relevant to your product. For each, draft two thresholds, one that requires named safeguards before release and one that stops further work until they exist, with the specific eval that would trigger each.</li>
96<li>Identify the safeguards required at the lower threshold. Decide which of those are already in place and which would need to be added before crossing.</li>
97</ol>
98<p>The exercise above is the smallest version of a Preparedness-style framework that is still a real commitment. It takes a senior engineering or product leader half a day to draft, plus a conversation with executive leadership to ratify the pause-or-restrict authority. The output is a one-page policy that closes a meaningful slice of audit and procurement questions and gives the eval team a concrete forward-looking <a href="/glossary#checklist" class="glossary-link" data-glossary-term="checklist">checklist</a>.</p>
99<p>The next chapter, <a href="/learn/governance/ai-risk-register">Building an AI risk register</a>, covers the artifact that complements the forward-looking policy with a current-state map.</p>
100<section data-footnotes="" class="footnotes"><h2 class="sr-only" id="footnote-label">Footnotes</h2>
101<ol>
102<li id="user-content-fn-openai-preparedness-v2">
103<p>OpenAI, "Preparedness Framework, Version 2." Last updated April 15, 2025. <a href="https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf" target="_blank" rel="noopener noreferrer">https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf</a> <a href="#user-content-fnref-openai-preparedness-v2" data-footnote-backref="" aria-label="Back to reference 1" class="data-footnote-backref">↩</a> <a href="#user-content-fnref-openai-preparedness-v2-2" data-footnote-backref="" aria-label="Back to reference 1-2" class="data-footnote-backref">↩<sup>2</sup></a> <a href="#user-content-fnref-openai-preparedness-v2-3" data-footnote-backref="" aria-label="Back to reference 1-3" class="data-footnote-backref">↩<sup>3</sup></a></p>
104</li>
105<li id="user-content-fn-openai-preparedness">
106<p>OpenAI, "Preparedness Framework." <a href="https://openai.com/safety/preparedness/" target="_blank" rel="noopener noreferrer">
106https://openai.com/safety/preparedness/</a> <a href="#user-content-fnref-openai-preparedness" data-footnote-backref="" aria-label="Back to reference 2" class="data-footnote-backref">↩</a> <a href="#user-content-fnref-openai-preparedness-2" data-footnote-backref="" aria-label="Back to reference 2-2" class="data-footnote-backref">↩<sup>2</sup></a></p>
107</li>
108<li id="user-content-fn-openai-path-to-astra">
109<p>OpenAI, "Path to Astra: critical capabilities and frontier safeguards." September 1, 2026. <a href="https://openai.com/index/path-to-astra/" target="_blank" rel="noopener noreferrer">https://openai.com/index/path-to-astra/</a> <a href="#user-content-fnref-openai-path-to-astra" data-footnote-backref="" aria-label="Back to reference 3" class="data-footnote-backref">↩</a></p>
110</li>
111<li id="user-content-fn-openai-pacing-cyber">
112<p>OpenAI, "Pacing model development in an era of cyber-critical capabilities." August 18, 2026. <a href="https://openai.com/index/pacing-model-development-cyber-capabilities/" target="_blank" rel="noopener noreferrer">
112https://openai.com/index/pacing-model-development-cyber-capabilities/</a> <a href="#user-content-fnref-openai-pacing-cyber" data-footnote-backref="" aria-label="Back to reference 4" class="data-footnote-backref">↩</a></p>
113</li>
114<li id="user-content-fn-anthropic-rsp">
115<p>Anthropic, "Responsible Scaling Policy." Version 3.4, effective July 8, 2026. <a href="https://www.anthropic.com/responsible-scaling-policy" target="_blank" rel="noopener noreferrer">https://www.anthropic.com/responsible-scaling-policy</a> <a href="#user-content-fnref-anthropic-rsp" data-footnote-backref="" aria-label="Back to reference 5" class="data-footnote-backref">↩</a> <a href="#user-content-fnref-anthropic-rsp-2" data-footnote-backref="" aria-label="Back to reference 5-2" class="data-footnote-backref">↩<sup>2</sup></a></p>
116</li>
117<li id="user-content-fn-openai-preparedness-repo">
118<p>OpenAI, "Frontier Evals" (formerly the preparedness repo: PaperBench, SWE-Lancer, EVMbench). <a href="https://github.com/openai/frontier-evals" target="_blank" rel="noopener noreferrer">https://github.com/openai/frontier-evals</a> <a href="#user-content-fnref-openai-preparedness-repo" data-footnote-backref="" aria-label="Back to reference 6" class="data-footnote-backref">↩</a></p>
119</li>
120<li id="user-content-fn-openai-safety-cases-training">
121<p>OpenAI, "Towards safety cases for frontier AI training." September 28, 2026. <a href="https://openai.com/index/towards-safety-cases-for-frontier-ai-training/" target="_blank" rel="noopener noreferrer">
121https://openai.com/index/towards-safety-cases-for-frontier-ai-training/</a> <a href="#user-content-fnref-openai-safety-cases-training" data-footnote-backref="" aria-label="Back to reference 7" class="data-footnote-backref">↩</a></p>
122</li>
123</ol>
124</section></article>`;export{e as default};
125//# sourceMappingURL=governance--openai-preparedness-C2_8ERUp.js.map

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.