Ashita Orbis
Reference

Preparedness, Responsible Scaling, and Frontier Safety: Comparing Lab Governance Frameworks

1. Why these frameworks matter

A Frontier AI Safety Framework is an attempt to turn a vague commitment—“do not deploy dangerously capable systems without adequate safeguards”—into a conditional operating procedure. The usual pattern is an “if-then commitment”: if a model reaches some capability threshold, then the developer must perform specified evaluations, apply stronger security or deployment mitigations, produce a risk assessment, obtain internal approval, disclose some information, or refrain from deployment until the risk is judged acceptable. The 2026 International AI Safety Report describes this as a prominent organizational approach to general-purpose AI risk management and notes that these frameworks commonly combine capability evaluations before mitigation with residual-risk analysis or safety-case-like reasoning after mitigation. International AI Safety Report

The three canonical lab frameworks use different labels for the same underlying governance problem. OpenAI uses “High” and “Critical” thresholds in Tracked Categories; Anthropic historically used AI Safety Levels (ASLs) and, in its 2026 rewrite, shifted toward capability thresholds plus arguments for safety; Google DeepMind uses Critical Capability Levels (CCLs) and, since April 2026, lower Tracked Capability Levels (TCLs). OpenAI’s latest public Preparedness Framework is Version 2, last updated April 15, 2025; Anthropic’s latest RSP is Version 3.2, effective April 29, 2026; Google DeepMind’s latest FSF is Version 3.1, published April 17, 2026. cdn.openai.com+2cdn.sanity.io+2

These frameworks are not regulation. They are voluntary corporate policies, and that distinction is central. They may be useful transitional artifacts: they define terminology, normalize pre-deployment dangerous-capability evaluations, create paper trails, and make future regulation easier to specify. But they also risk becoming a substitute for external enforcement: the same firms that profit from deployment often retain final authority over thresholds, acceptable residual risk, redactions, and whether competitors’ actions justify weaker safeguards. International AI Safety Report

2. The canonical frameworks at a glance

Lab Latest public framework covered here Threshold vocabulary Primary risk domains Basic gating claim
OpenAI Preparedness Framework v2, April 15, 2025 High and Critical capability thresholds inside Tracked Categories Biological/chemical capabilities, cybersecurity, AI self-improvement Do not deploy systems that reach High until associated severe-harm risks are sufficiently minimized; Critical capabilities require safeguards during development regardless of deployment plans. cdn.openai.com
Anthropic Responsible Scaling Policy v3.2, effective April 29, 2026 Historically ASL-1 through ASL-4+; latest version uses capability or usage thresholds plus safety arguments, while ASLs remain a label for present mitigation levels Chemical/biological weapons, high-stakes sabotage, automated R&D in key domains, with emphasis on AI R&D Publishes company plans, industry-wide recommendations, Risk Reports, external-review procedures, and competitor-contingent commitments; no longer presents future safeguards primarily as rigid ASL control lists. cdn.sanity.io+2cdn.sanity.io+2
Google DeepMind Frontier Safety Framework v3.1, April 17, 2026 Critical Capability Levels and Tracked Capability Levels CBRN, cyber, harmful manipulation, ML R&D and misalignment Evaluate proximity to T/CCLs across the model lifecycle; if thresholds are reached, apply security and deployment mitigations and make residual-risk determinations, with safety cases for CCLs. storage.googleapis.com+2storage.googleapis.com+2

A useful mental model is that all three frameworks are variants of Capability Threshold Governance. They do not primarily ask, “Is this model good or bad?” They ask a more operational question: “Has the model crossed a capability line such that ordinary release processes are no longer enough?” That makes the central epistemic problem unavoidable: the frameworks only work if the lab can identify the relevant capability, elicit it under realistic adversarial conditions, map it to plausible harm, and assess whether mitigations actually reduce residual risk.

3. OpenAI’s Preparedness Framework

OpenAI’s Preparedness Framework is the most explicitly board-facing of the three canonical documents. Version 2 defines its purpose as tracking and preparing for frontier capabilities that create new risks of “severe harm,” which OpenAI defines as death or grave injury to thousands of people or hundreds of billions of dollars in economic damage. It currently focuses on three Tracked Categories: biological and chemical capabilities, cybersecurity capabilities, and AI self-improvement capabilities. cdn.openai.com

3.1 Capability tiers: High and Critical

The key governance distinction is between High and Critical capability thresholds. High means the model significantly increases existing risk vectors for severe harm; covered systems crossing High require robust safeguards before deployment and appropriate security controls during development. Critical means the model presents a meaningful risk of a qualitatively new severe-harm threat vector with no ready precedent; Critical capabilities require safeguards during development regardless of deployment plans. cdn.openai.com

OpenAI’s v2 change log says it removed “low” and “medium” terminology because those levels were not operationally involved in Preparedness work. It also moved persuasion outside the Preparedness Framework, placed nuclear and radiological capabilities into Research Categories, and focused Tracked Categories on biological/chemical, cyber, and AI self-improvement risks. cdn.openai.com

This matters because OpenAI’s framework is deliberately narrow. It is not a general AI ethics policy, a broad social-harms framework, or a comprehensive safety-management system. It is a catastrophic-risk governance framework for a small number of capability classes that OpenAI judges plausible, measurable, severe, net-new, and instantaneous or irremediable. That focus makes the framework tractable, but it also means many important harms sit outside it by design. cdn.openai.com

3.2 Evaluation and elicitation

OpenAI emphasizes that dangerous-capability evaluations should approximate what an adversary could extract from the model, including through high-capability settings, variants with negligible refusal rates on evaluation tasks, and the best available scaffolds. It explicitly states that one-time elicitation is a lower bound rather than a ceiling, because scaffolding and elicitation methods continue to improve after evaluation. cdn.openai.com

That admission is important. A model can be below threshold under today’s elicitation stack and above threshold under tomorrow’s agent scaffold, fine-tuning recipe, tool-use loop, or jailbreak. The framework’s scientific burden is therefore not merely benchmark design; it is forecasting adversarial elicitation progress. This is one form of the Threshold-Setting Problem.

OpenAI divides evaluations into scalable evaluations and deep dives. Deep dives may include human expert red-teaming, expert consultation, resource-intensive third-party evaluations such as biological wet-lab studies, and independent third-party evaluator assessments. cdn.openai.com

3.3 Safeguards, residual risk, and deployment gates

For systems near or above a capability threshold, OpenAI compiles planned safeguards into a Safeguards Report. The report is supposed to map severe-harm pathways to security controls and safeguards, describe safeguard efficacy, assess residual risk, and note limitations. The Safety Advisory Group (SAG) assesses whether safeguards sufficiently minimize risk, then recommends deployment, further evaluation, or alternative deployment conditions; final decisions go to OpenAI Leadership. cdn.openai.com

OpenAI’s safeguard taxonomy distinguishes risks from malicious users and risks from misaligned models. For malicious users, potential claims include robust refusal behavior, usage monitoring, and trust-based access. For misaligned models, potential claims include lack of autonomous capability, value alignment, instruction alignment, robust oversight, and restricted system architecture. cdn.openai.com

3.4 Governance: SAG, leadership, and the board

OpenAI’s governance structure gives the SAG responsibility for reviewing reports, assessing capability levels and residual risks, and recommending next steps. But the document is explicit that the SAG cannot “filibuster”: OpenAI Leadership can make decisions without SAG participation, and the CEO or a CEO-designated person is responsible for final decisions, including accepting residual risks and making deployment go/no-go calls. The Board’s Safety and Security Committee receives visibility, can review decisions, can require reports and information, and may reverse a decision or mandate a revised course of action. cdn.openai.com

This is stronger than a purely advisory ethics review, because the board committee has an explicit reversal power. It is weaker than external regulation, because the final operational gate remains inside OpenAI unless law, contract, or some other external mechanism constrains the decision.

3.5 Transparency and third-party participation

OpenAI commits to public disclosures for major deployments, including testing scope, capability evaluations for each Tracked Category, reasoning for deployment, and decisive context about model development or capabilities. If a model is beyond a High threshold, OpenAI says it will include information about safeguards, subject to redaction or summarization for reasons such as intellectual property or safety. OpenAI also says it will work with third parties for independent capability evaluation or safeguard stress-testing when it deems deeper testing warranted and when high-quality external testing is available. cdn.openai.com

The transparency commitment is meaningful but conditional. The company decides what warrants deeper testing, when external testing is feasible, and what must be redacted. The framework creates an expectation of third-party participation, not an unconditional right of third-party inspection.

3.6 Marginal risk and competitor behavior

OpenAI’s marginal-risk clause is one of the most important and controversial structural features. If another frontier developer releases a High or Critical system without comparable safeguards, OpenAI says it could adjust its own safeguard requirements, but only if doing so does not meaningfully increase overall severe-harm risk, if it publicly acknowledges the adjustment, and if it keeps safeguards more protective than the other developer while sharing information to validate that claim. cdn.openai.com

The strongest reading is pragmatic: if the baseline risk has already changed because another model is available, absolute risk reduction from unilateral restraint may be limited. The weaker reading is that this creates a race-to-the-bottom escape hatch: once one actor lowers standards, others can describe their own weaker posture as low marginal risk.

4. Anthropic’s Responsible Scaling Policy

Anthropic’s Responsible Scaling Policy is the framework most associated with the term “responsible scaling.” It is also the framework whose 2026 revisions most visibly expose the difficulty of unilateral self-regulation in a competitive frontier-model market. Version 3.2 describes the RSP as Anthropic’s voluntary framework for managing catastrophic risks from advanced AI systems and says it establishes how Anthropic identifies and evaluates risks, makes development and deployment decisions, and aims to ensure that model benefits exceed costs. cdn.sanity.io

4.1 The ASL lineage

The original RSP vocabulary was organized around AI Safety Levels, by analogy to biosafety levels. The 2026 International AI Safety Report, summarizing Anthropic’s earlier RSP 2.2, describes ASL-1 as no significant catastrophic risk, ASL-2 as early signs of dangerous capabilities requiring ASL-2 deployment and security standards, ASL-3 as substantially increased catastrophic misuse risk requiring ASL-3 deployment and/or security standards, and ASL-4+ as future classifications not yet defined. International AI Safety Report

By v3.2, however, Anthropic no longer treats future capability levels as primarily a rigid ASL ladder. Appendix B says earlier editions defined ASLs with specific required controls; Anthropic still uses ASLs to refer to present levels of risk mitigations for existing models, but for future capabilities it prefers to focus on what sort of argument a developer should make about risk containment. cdn.sanity.io

This shift is philosophically important. The framework moves from “capability level X requires control bundle Y” toward “capability threshold X requires a strong safety argument addressing specified actors and pathways.” That is closer to Safety Case regulation, but it is also less mechanically enforceable unless an external party can judge the argument.

4.2 Version 3: the collective-action rewrite

Anthropic says the third iteration changed because of a collective-action problem: societal catastrophic risk depends on multiple AI developers, not just Anthropic. If one developer pauses while others proceed without strong mitigations, Anthropic argues the result could be less safe because weaker-protection developers would set the pace and more safety-focused developers would lose influence and safety-research capacity. cdn.sanity.io

The new structure separates “our plans as a company” from “industry-wide recommendations.” Anthropic explicitly says it cannot unilaterally and unconditionally commit to staying aligned with those industry-wide recommendations, though it uses them as a north star for mitigation planning and public-policy work. It also adopts competitor-contingent commitments for scenarios where it can be confident other relevant developers are acting similarly. cdn.sanity.io

This is the clearest internal acknowledgment, among the three frameworks, that voluntary lab governance is structurally unstable under competition. Anthropic’s RSP v3.2 is not merely a safety policy; it is a diagnosis of why one-lab safety policies may fail without industry-wide governance.

4.3 Capability thresholds and required protections

Anthropic’s v3.2 table maps capability or usage thresholds to company plans and industry-wide recommendations. For non-novel chemical/biological weapons production, the threshold is an AI system that significantly helps individuals or groups with basic technical backgrounds create, obtain, and deploy chemical or biological weapons with serious catastrophic potential. Anthropic’s company plan is to maintain or improve ASL-3 protections, including robust classifier guards, access controls for trusted users, red-teaming, bug bounties, threat intelligence, and security controls. cdn.sanity.io

For novel chemical/biological weapons production, the relevant actors become better-resourced threat actors, and Anthropic says the required argument and protections should rise correspondingly, likely including security roughly in line with RAND Security Level 4, depending on the strongest plausible threat actors not bound by a credible governance regime. cdn.sanity.io

For high-stakes sabotage, Anthropic focuses on AI systems that are highly relied on, have extensive access to sensitive assets, and have moderate autonomous, goal-directed, and subterfuge-relevant capabilities. The industry-wide safety argument would need evidence such as compartmentalization, code review, capability assessments showing inability to carry out irreversible sabotage, monitoring or restriction of internal AI behavior, and evidence that models lack propensities to deceive, manipulate, or sabotage users. cdn.sanity.io

For automated R&D in key domains, Anthropic defines the threshold as systems that can fully automate or dramatically accelerate large, top-tier research teams in areas such as energy, robotics, weapons development, and AI itself. For now, Anthropic says its evaluations focus on AI R&D because that domain is more tractable and closer to current model strengths; it considers the threshold met if models can substitute for its full set of research scientists and engineers at competitive costs within a factor of five, or if AI progress dramatically accelerates in a way plausibly attributable to AI R&D automation. cdn.sanity.io

4.4 Risk Reports and external review

Anthropic’s v3.2 transparency machinery is more elaborate than OpenAI’s. It commits to publishing public Risk Reports, minimizing redactions, and seeking comprehensive public external review of Risk Reports. External reviewers are to receive private versions with no or minimal redactions, publish commentary on reasoning quality, risk-assessment validity, overall risk level, and appropriateness of redactions, and be chosen with attention to expertise, candor incentives, and conflicts of interest. cdn.sanity.io

The document is unusually candid that there are no well-established organizations or procedures for this kind of external review and that Anthropic is treating the process as an experiment. That candor is valuable; it also means the mechanism should not be mistaken for a mature audit regime. cdn.sanity.io

4.5 Governance: CEO, RSO, Board, and LTBT

Anthropic’s Risk Report process sends the report, internal feedback, and any available external feedback to the CEO and Responsible Scaling Officer (RSO), who make the ultimate determination about the adequacy of the risk assessment and downstream development or deployment plans. The Board and Long Term Benefit Trust (LTBT) receive the decision, report, and internal feedback. If marginal-risk analysis plays a major role in moving forward, explicit Board and LTBT approval is required. cdn.sanity.io

Anthropic also commits to maintaining an RSO, briefing the LTBT, sharing unredacted final Risk Reports with regular-clearance staff, allowing anonymous or identified noncompliance reports, protecting reporters from retaliation, avoiding safety-impeding non-disparagement obligations, conducting internal review, commissioning annual third-party procedural compliance review, and requiring Board approval in consultation with the LTBT for policy changes. cdn.sanity.io

This is the strongest internal-accountability package among the three frameworks. Its main weakness is that much of the strongest oversight remains procedural rather than substantive: a third party may assess whether Anthropic followed the policy, but not necessarily whether the policy’s risk judgments were correct.

4.6 Competitor-contingent commitments

Anthropic’s Appendix A makes competitor contingency explicit. If Anthropic is in the lead and has clear evidence no competitor will soon develop a similarly highly capable model, it says it will require a strong argument that catastrophic risk is contained and delay development or deployment as needed. If competitors have strong safety measures, Anthropic says it will meet or exceed their overall risk-reduction posture and delay as needed until it can. If a competitor implements a significantly better mitigation, Anthropic says it will make a significant effort to meet or exceed that standard, though it will not necessarily delay development or deployment in that scenario. cdn.sanity.io

This is simultaneously a safety provision and a race-dynamics provision. It tries to avoid unilateral disadvantage while preserving some duty to match or exceed the safety posture of peers. The hard question is whether such clauses stabilize cooperation or make standards endogenous to the least cautious credible competitor.

5. Google DeepMind’s Frontier Safety Framework

Google DeepMind’s Frontier Safety Framework is the most formal in its terminology of capability levels and lifecycle risk-management process. Version 3.1 defines the framework as a set of protocols for severe risks arising from high-impact frontier AI capabilities. Its core components are identifying capability levels that could pose severe risk, detecting those levels across the lifecycle, preparing mitigation plans, and involving external parties where required or appropriate. storage.googleapis.com

5.1 CCLs and TCLs

The framework is built primarily around Critical Capability Levels. A CCL is a capability level at which, absent mitigation, a frontier AI model or system may pose heightened risk of severe harm. CCLs are derived by identifying foreseeable pathways through which a model could cause severe harm and defining the minimal capabilities required for that pathway. storage.googleapis.com

Version 3.1 adds Tracked Capability Levels, which are lower thresholds meant to capture significant, but not severe, risks. Google DeepMind identifies TCLs for CBRN risk and for ML R&D/misalignment risk, and says early-warning evaluations will be used to assess proximity to TCLs and CCLs, involving internal and external experts as needed. storage.googleapis.com

The risk domains are CBRN, cyber, harmful manipulation, and ML R&D/misalignment. The inclusion of harmful manipulation distinguishes Google DeepMind from OpenAI’s narrower Preparedness v2 scope, where persuasion is handled outside the Preparedness Framework. storage.googleapis.com

5.2 Lifecycle risk management

Google DeepMind says it conducts risk management throughout model development, on checkpoints and versions, before and after deployment. It performs critical capability assessments before first external deployment of a new frontier model and performs material capability change assessments for subsequent versions when meaningful capability increases might undermine prior risk justifications. storage.googleapis.com

When a model reaches an alert threshold for a CCL, Google DeepMind assesses proximity to the CCL, analyzes risk, involves internal and external experts as needed, and formulates a response plan. The two main mitigation categories are security mitigations, such as preventing model-weight exfiltration, and deployment mitigations, such as safety fine-tuning, monitoring, and response. storage.googleapis.com

For models reaching TCLs or CCLs, Google DeepMind performs residual-risk assessments. For models reaching CCLs, that residual-risk assessment is informed by a supplemental safety case. storage.googleapis.com

5.3 Risk acceptance and marginal risk

Google DeepMind’s risk acceptance approach is explicitly residual-risk-based. If no T/CCL is reached, the model is deemed acceptable for further development and deployment because it should not possess the relevant capabilities. If a misuse T/CCL is reached, deployment can still be acceptable if mitigations reduce residual harm to an acceptable level and security mitigations are adequate. The framework explicitly considers the capabilities and mitigations of other publicly available models, including the possibility that if similarly capable models with fewer mitigations are already public, the marginal risk added by deployment may be low. storage.googleapis.com

For ML R&D CCLs, Google DeepMind recommends Security Level 3 for one acceleration threshold and Security Level 4 for full automation of a Google AI-capabilities research team, while emphasizing that the stronger level must be taken on by the frontier AI field as a whole. storage.googleapis.com

As with OpenAI and Anthropic, the marginal-risk logic is double-edged. It is rational under non-ideal competition, but dangerous if it becomes a general license to weaken controls whenever the market baseline worsens.

5.4 Governance and disclosure

Google DeepMind’s governance language is less specific than OpenAI’s board-reversal structure or Anthropic’s RSO/LTBT machinery. It says Google has a comprehensive internal governance structure with clearly allocated responsibilities, including legal, compliance, and safety reviews with escalation procedures. It also commits to reviewing the framework at least annually, or more often if its adequacy or adherence is materially undermined. storage.googleapis.com

On disclosure, Google DeepMind says that if a model reaches a CCL that poses unmitigated and material public-safety risk, it aims to share relevant information with appropriate government authorities, including model information, evaluation results, and mitigation plans where appropriate and subject to confidentiality, security, proprietary, and sensitive-information constraints. It may also consider disclosure to other external organizations. storage.googleapis.com

This is a weaker public-transparency commitment than Anthropic’s Risk Report model. It is more government-oriented and less structured around public external review.

6. Structural similarities

6.1 Capability thresholds gate deployment and development

All three frameworks use capability thresholds as governance triggers. This is the core similarity. OpenAI gates deployment at High and development safeguards at Critical. Anthropic maps capability or usage thresholds to company plans and industry-wide safety arguments. Google DeepMind maps T/CCLs to mitigation and residual-risk processes. cdn.openai.com+2cdn.sanity.io+2

This makes them all instances of Threshold-Based AI Governance. The threshold is not merely a benchmark score; it is a claim about a model’s ability to contribute to a harm pathway. That is why all three frameworks rely on threat modeling, expert judgment, red-teaming, and post-mitigation residual-risk assessment.

6.2 The frameworks are safety-case-like

All three are moving toward Safety Case reasoning. OpenAI’s Safeguards Reports compile evidence that safeguards sufficiently minimize severe-harm risk. Anthropic’s v3.2 explicitly structures industry-wide recommendations around strong arguments for safety rather than fixed ASL control lists. Google DeepMind says residual-risk assessments for models reaching CCLs are informed by supplemental safety cases. cdn.openai.com+2cdn.sanity.io+2

This is a healthy development insofar as benchmark scores alone cannot establish safety. It is also a source of discretion. A safety case is only as good as the independence, competence, and authority of the institution judging it.

6.3 Security and deployment mitigations are both central

The frameworks all distinguish between preventing misuse through access channels and preventing dangerous access to the model itself. OpenAI discusses safeguards against malicious users, safeguards against misaligned models, and security controls. Anthropic’s ASL-3 protections include classifiers, access controls, red-teaming, bug bounties, threat intelligence, and security controls. Google DeepMind explicitly divides mitigations into security mitigations, such as preventing model-weight exfiltration, and deployment mitigations, such as safety fine-tuning and monitoring. cdn.openai.com+2cdn.sanity.io+2

The shared assumption is defense-in-depth: no single layer—refusal training, monitoring, access control, model-weight security, red-team testing, or legal policy—is sufficient by itself. The 2026 International AI Safety Report describes defense-in-depth for general-purpose AI as combining technical, organizational, and societal measures across development and deployment. International AI Safety Report

6.4 Third-party evaluation is expected, but conditional

All three frameworks gesture toward third-party evaluation or external input. OpenAI says it will work with third parties for independent capability evaluation and safeguard stress-testing when deeper testing is warranted and high-quality testing is available. Anthropic commits to external review of Risk Reports under specified circumstances and gives the LTBT authority to request external review. Google DeepMind says it may involve external parties where required or appropriate and may engage governments or other external actors to inform responsible development and deployment. storage.googleapis.com+3cdn.openai.com+3cdn.sanity.io+3

The key caveat is that none of these frameworks creates a general public inspection right. Third-party evaluation is mediated by the company, limited by confidentiality and safety redactions, and dependent on a still-immature evaluator ecosystem.

6.5 Competition is now inside the frameworks

All three frameworks account for competitor behavior. OpenAI’s marginal-risk clause allows adjustment if another developer releases High or Critical capabilities without comparable safeguards. Anthropic’s v3.2 separates unilateral company plans from industry-wide recommendations and uses competitor-contingent commitments. Google DeepMind considers other publicly available models and their mitigations in assessing marginal deployment risk. storage.googleapis.com+3cdn.openai.com+3cdn.sanity.io+3

This is probably inevitable. Frontier AI safety is not a single-firm control problem; it is a strategic interaction among labs, governments, users, cloud providers, open-weight ecosystems, and potential attackers. But embedding competition inside safety commitments makes the commitments less stable.

7. Structural differences

Dimension OpenAI Preparedness Framework Anthropic RSP Google DeepMind FSF
Governance specificity Specific SAG process; final decisions by CEO or delegate; Board Safety and Security Committee can reverse decisions. cdn.openai.com CEO and RSO approve Risk Reports; Board and LTBT receive reports; Board/LTBT approval required when marginal-risk analysis is central; explicit noncompliance and employee-speech protections. cdn.sanity.io Describes comprehensive internal governance, legal/compliance/safety reviews, and escalation, but provides less public detail about named decision bodies and veto authority. storage.googleapis.com
Transparency Public deployment disclosures for major deployments, with redactions; third-party evaluation conditional. cdn.openai.com Public Risk Reports, redaction minimization, external review process, reviewer conflict criteria, and public commentary. cdn.sanity.io Government-oriented disclosure for unmitigated material public-safety CCL risk; may disclose to other external organizations. storage.googleapis.com
Reversibility / pausing Do not deploy High until risk minimized; Critical requires safeguards during development; board may reverse decisions. cdn.openai.com Will delay in some lead or peer-safety scenarios; says it may pause even outside listed commitments, but the structure is explicitly competitor-contingent. cdn.sanity.io Uses risk acceptance and response plans; less explicit public language about hard pauses or deletion-style reversibility. storage.googleapis.com
Treatment of future uncertainty Treats one-time elicitation as a lower bound and updates evaluations as scaffolding changes. cdn.openai.com Says it cannot give highly specific advance detail on future evaluations or mitigations; latest version favors safety arguments over rigid ASL lists. cdn.sanity.io Says AI risk assessment science is still developing and assessments will often involve subjective analysis. storage.googleapis.com
Scope Narrow catastrophic-risk categories; persuasion and nuclear/radiological are outside Tracked Categories in v2. cdn.openai.com Catastrophic risks, with focus on CBRN, sabotage, and automated R&D; not comprehensive for all AI obligations. cdn.sanity.io CBRN, cyber, harmful manipulation, ML R&D, and misalignment; harmful manipulation is inside the framework. storage.googleapis.com

The headline difference is not that one lab has thresholds and another does not. They all have thresholds. The difference is where discretion sits after the threshold is triggered.

OpenAI’s framework is operationally concrete about internal decision flow but leaves final deployment authority with leadership, subject to board oversight. Anthropic’s framework is the most explicit about public reporting, external review, and internal accountability, but its 2026 rewrite weakens the idea of unconditional unilateral commitments. Google DeepMind’s framework has the most systematic CCL/TCL vocabulary and lifecycle risk-management structure, but its public governance and disclosure commitments are less detailed than Anthropic’s and less board-specific than OpenAI’s. cdn.openai.com+2cdn.sanity.io+2

8. Active critiques

8.1 Bengio’s self-regulation critique

Yoshua Bengio’s critique is best read as institutional rather than merely technical. In July 2024, he argued that conflicts of interest between public good and profit maximization inside corporate AI labs may require governments to force broader stakeholder representation, and he posed the release/open-source line as a question for democratically chosen governments rather than CEOs. In a later public statement, he argued that AI companies need strong regulation, clear red lines, substantial penalties, and enforcement mechanisms, and that self-regulation is not the answer for the industry. Yoshua Bengio

Applied to RSPs and adjacent frameworks, the critique is straightforward: a voluntary scaling policy can be well-intentioned and still structurally inadequate. If the company defines the thresholds, controls the evidence, selects or funds evaluators, redacts the public report, and retains final authority to proceed, then the policy is not a substitute for external governance. It is corporate self-regulation with some transparency and process discipline.

Anthropic’s v3.2 partly concedes the point. It says the best way for its industry-wide recommendations to be implemented is likely through governance of all relevant frontier developers by third parties that determine who must provide risk analyses and whether those arguments are adequate; it also says national regulation should be harmonized to avoid a race to the bottom. cdn.sanity.io

8.2 Lack of teeth without external enforcement

The strongest critique is not that the frameworks are useless. It is that their practical force depends on implementation by the same organizations whose incentives are being constrained. The 2026 International AI Safety Report says external compliance assessments remain limited because frameworks are recent, public information is scarce, and there are no standardized external audits; it also says the frameworks may not ensure effective risk management on their own and vary substantially in scope, thresholds, and enforceability. International AI Safety Report

A 2025 affordance analysis of OpenAI’s Preparedness Framework v2 argues that prominent AI safety frameworks function as voluntary self-governance and that effective mitigation requires more robust governance interventions beyond current industry self-regulation. The paper’s method is notable because it asks what actions a policy actually permits, demands, discourages, or refuses—not merely what risks it rhetorically names. arXiv

SaferAI’s April 2026 comparative evaluation similarly argues that providers retain discretion at key decision points, that final authority often rests with executive leadership, and that discretionary language and deferred mitigation specification limit external scrutiny. It also warns that competitor-relative standards can create interdependencies in which one provider’s weaker standards lower the baseline for others. arXiv

8.3 The threshold-setting epistemic problem

The frameworks depend on thresholds, but threshold-setting is scientifically immature. OpenAI says one-time elicitation is only a lower bound because scaffolding and elicitation methods improve. Anthropic says it cannot presently give highly specific advance detail on future evaluations or mitigations and that the science of evaluations is not mature enough to confidently predict the precise buffer between current models and a capability threshold. Google DeepMind says frontier AI risk assessment is complex, the science is still developing, and assessments will often involve subjective analysis. storage.googleapis.com+3cdn.openai.com+3cdn.sanity.io+3

This problem has several components. First, dangerous capabilities may be latent: a model may possess them but not reveal them under the evaluation scaffold. Second, model access conditions matter: API-only access, tool access, fine-tuning, agent scaffolding, and open weights create different threat models. Third, capability thresholds are not equivalent to harm thresholds: a model’s ability to help design a pathogen is not the same as an actor’s ability to acquire materials, run a lab, evade detection, and release it. Fourth, post-deployment learning by users can change effective capability even if model weights do not change. Fifth, deceptive or sandbagging behavior would undermine evaluations precisely when evaluation reliability matters most.

The threshold problem is not a reason to abandon thresholds. Without thresholds, governance becomes pure discretion. But it is a reason to treat thresholds as defeasible risk indicators, not bright metaphysical lines between safe and unsafe systems.

8.4 External evaluation and audit are underdeveloped

Third-party evaluation is often invoked as the cure for self-regulation, but the current ecosystem is not yet mature. Anthropic explicitly says there are no well-established organizations or procedures for comprehensive public external review of Risk Reports. SaferAI’s 2026 comparison finds that no assessed provider commits to non-interference with external evaluation findings, that proof of containment or deployment-measure sufficiency is weak, and that third-party verification of deployment measures is almost absent. cdn.sanity.io

The problem is not merely institutional capacity. It is also access. A serious evaluator may need model weights, fine-tuning access, scaffolding access, logs, internal incident reports, cyber-security architecture, model-card drafts, evaluation failures, and deployment telemetry. Much of that information is commercially sensitive, security-sensitive, or legally constrained. A credible audit regime must solve the information-access problem without turning auditors into either rubber stamps or leak risks.

8.5 Scope gaps

The frameworks focus on a subset of risks. OpenAI v2 excludes persuasion from the Preparedness Framework and moves nuclear/radiological capabilities into Research Categories. Google DeepMind includes harmful manipulation, but its disclosure is limited to CCLs posing unmitigated and material public-safety risk. Anthropic says its RSP focuses on catastrophic risks and is not comprehensive for all obligations. cdn.openai.com+2storage.googleapis.com+2

The 2026 International AI Safety Report notes that frontier safety frameworks tend to focus on a subset of risk domains and that some prominent risks receive less emphasis. It also notes that, unlike risk management in sectors such as aviation or nuclear power, these frameworks typically do not use explicit quantitative risk thresholds. International AI Safety Report

This creates a classification problem: if a risk is excluded because it is genuinely lower severity, the framework remains focused. If it is excluded because it is hard to measure, socially diffuse, politically inconvenient, or not easily expressible as a capability threshold, the framework under-governs important harms.

9. Are voluntary frameworks stable, or a transition pattern toward regulation?

There are two plausible interpretations.

The optimistic interpretation is that voluntary frameworks are a transition pattern. They create the vocabulary regulators need: capability thresholds, dangerous-capability evaluations, residual-risk assessments, safety cases, deployment mitigations, model-weight security levels, Risk Reports, external review, incident reporting, and governance roles. The 2026 International AI Safety Report notes that early regulatory approaches are beginning to introduce legal requirements for standardization and transparency, including the EU AI Act, the EU General-Purpose AI Code of Practice, South Korea’s AI framework law, and California SB 53’s transparency requirements for safety frameworks and incident reporting. International AI Safety Report

The pessimistic interpretation is that voluntary frameworks become a liability shield and a political substitute for enforceable rules. A lab can publish a framework, make conditional commitments, disclose partial information, and still retain discretion over the central questions: whether a threshold was crossed, whether risk is acceptable, whether redactions are material, whether third-party evaluation is warranted, and whether competitors’ behavior justifies moving faster.

The most likely near-term path is hybridization. Voluntary frameworks will not disappear; they will be incorporated into regulatory, procurement, insurance, and investor expectations. But their stabilizing power will depend on whether external institutions can impose five things: independent access to evidence, enforceable red lines, standardized reporting, penalties for misrepresentation or noncompliance, and protection for internal dissenters who surface safety-relevant information.

Anthropic’s v3.2 effectively says this out loud: one developer cannot unilaterally guarantee industry-wide safety, and the best implementation of its recommendations likely involves third-party governance of relevant frontier developers. Google DeepMind similarly says some mitigations have social value only if broadly adopted across industry. OpenAI’s marginal-risk clause also assumes that competitor behavior changes the risk baseline. These are not incidental footnotes; they are admissions that frontier safety is a collective-action problem. cdn.sanity.io+2storage.googleapis.com+2

10. Bottom line

OpenAI’s Preparedness Framework, Anthropic’s Responsible Scaling Policy, and Google DeepMind’s Frontier Safety Framework converge on the same architectural idea: dangerous capabilities should trigger stronger evaluations, safeguards, security controls, residual-risk analysis, and governance review before deployment or further scaling. This convergence is real progress. It gives labs, governments, auditors, and civil society a shared object to critique.

But the convergence should not be mistaken for sufficient governance. The frameworks remain voluntary, threshold science is immature, third-party evaluation is conditional and underdeveloped, and competitor-relative clauses introduce strategic fragility. The unresolved governance question is whether these frameworks become enforceable scaffolding for public oversight or remain sophisticated self-regulatory documents whose hardest decisions are still made inside the labs.

Companion entries

Core theory: Frontier AI Safety Frameworks, Capability Threshold Governance, Safety Cases, Responsible Scaling, Catastrophic AI Risk, Defense in Depth

Canonical lab frameworks: OpenAI Preparedness Framework, Anthropic Responsible Scaling Policy, Google DeepMind Frontier Safety Framework, AI Safety Levels, Critical Capability Levels, Tracked Capability Levels

Evaluation and evidence: Dangerous Capability Evaluations, Capability Elicitation, Scaffolding and Tool Use, Third-Party Model Evaluation, AI Red-Teaming, AI Sandbagging, Residual Risk Assessment

Governance and institutions: Board-Level AI Safety Governance, AI Auditing, Frontier Model Reporting, External Review of AI Systems, Whistleblower Protection in AI Labs, Safety Case Regulation

Critiques and open problems: Self-Regulation Critique, Threshold-Setting Problem, Marginal Risk Arguments, Race to the Bottom in AI Safety, Voluntary AI Commitments, Reversibility in AI Deployment

AI-researched reference article. Follow the citations for load-bearing claims; corrections welcome via contact.