Tracked Capability Levels: Governance Through Threshold Tiering
Three frontier AI developers — Anthropic, OpenAI, and Google DeepMind — have converged on a similar governance pattern: define discrete capability tiers, evaluate models against them, and commit to specific safeguard or deployment actions when a tier is reached. This article describes the operational concept, maps the three published frameworks, catalogs the documented tier-related events through mid-2026, and treats the open question of whether threshold tiering is a stable governance pattern or transitional scaffolding for external regulation.
Coverage note: verified through May 10, 2026.
The operational concept
Threshold tiering is governance by tripwire. Each lab publishes a small number of capability levels, each defined to roughly correspond to a category of harm or risk that the lab considers worth treating differently. Each level is paired with at least three things: an evaluation regime that determines whether a model has reached the level, a set of safeguards or governance actions that activate once the level is reached, and a decision authority responsible for declaring the level reached or not.
The mechanism is intentionally coarse. A continuous risk surface — say, a model's contribution to biological-weapons uplift or its capacity to autonomously execute cyber operations against hardened targets — is collapsed into a small number of bins. The bins exist because regulators, internal review boards, and partner organizations need a small number of categorical decisions, not a continuous score. A model is either approved for deployment, approved with controls, restricted to limited internal use, or paused entirely. Tiering supplies the labels those decisions hang on. See Precautionary Reasoning Under Capability Uncertainty for the broader framing.
The conceptual heritage is two-fold. The Anthropic Responsible Scaling Policy explicitly draws an analogy to the U.S. government's Biosafety Level system, in which physical containment standards escalate with pathogen severity (Anthropic RSP v3.0). The OpenAI Preparedness Framework draws on cybersecurity vulnerability triage and defense-acquisition risk classification, in which threats are sorted into tiers tied to specific response procedures (OpenAI Preparedness Framework v2). Google DeepMind's Frontier Safety Framework treats Critical Capability Levels as analogous to dangerous-goods classification, with internal handling and external deployment controls increasing in step with hazard category (Google DeepMind FSF 3.1).
A subtler property of these frameworks is that the tiers are not directly observable. A tier is a determination based on evaluation evidence, plus interpretation, plus a judgment call about whether elicitation was adequate. The tier is the output of a process, not a measurement. This matters because it shapes both the strengths of the system — labs can incorporate qualitative judgment, classified threat models, and external eval input — and the weaknesses, which include the absence of a verifiable external referent against which the determination can be checked. The point is developed at length in Capability Elicitation and the Lower-Bound Problem.
Three published implementations
Anthropic: AI Safety Levels and the Responsible Scaling Policy
Anthropic publishes a Responsible Scaling Policy (RSP) that defines AI Safety Levels (ASLs) and conditional commitments triggered by them. The policy has gone through several revisions; the current version, RSP v3.0, was published in February 2026 and represents a substantial rewrite of the original 2023 scheme (Anthropic RSP v3.0).
The RSP commits Anthropic to running capability evaluations on its frontier models against a defined set of capability thresholds. When a threshold is reached or cannot be ruled out, Anthropic activates additional safeguards — model deployment mitigations and internal security controls — and produces a public document describing the assessment and the resulting decision. The CEO and the Responsible Scaling Officer hold ultimate decision authority over whether risk assessments are adequate and whether deployment plans satisfy the policy's commitments. An external review of procedural compliance is to occur at least annually.
RSP v3.0 makes two changes worth noting. First, it separates Anthropic's company-specific commitments from a set of broader recommendations the company believes industry should adopt. The policy explicitly states that managing ecosystem-wide risk depends on multiple developers, and that if Anthropic alone pauses while others continue, weaker-protection developers may set the pace. This is an admission that unilateral restraint is structurally unstable. Second, the policy refines the older binary "ASL-3 / not ASL-3" framing into capability thresholds tied to specific risk domains, with separate determinations for chemical, biological, radiological, and nuclear (CBRN) uplift, autonomous AI research and development, and other domains as they mature.
The first activation of ASL-3 protections occurred with Claude Opus 4 in May 2025 (Anthropic: Activating ASL-3 protections). Anthropic stated that it could not rule out ASL-3 capability in CBRN uplift and applied the corresponding deployment and security safeguards as a precaution. The activation was explicitly framed as not implying definitive evidence that the model had crossed the threshold; it was a procedural response to uncertainty. See ASL-3 Safeguards: Anthropic's First Activation for the operational details of what those safeguards include.
OpenAI: the Preparedness Framework
OpenAI's Preparedness Framework, currently at version 2, defines categories of "tracked risk" — biological and chemical, cybersecurity, AI self-improvement — and assigns models to capability levels within each: Low, Medium, High, and Critical, depending on the version (OpenAI Preparedness Framework v2, OpenAI: Updating our Preparedness Framework). The 2025 update removed Low and Medium from public reporting in some categories and reorganized persuasion as a research category rather than a tracked risk, which is itself a small example of how mutable these frameworks remain.
The Preparedness Framework ties High capability to pre-deployment safeguard requirements: a model classified High in a tracked risk category may not be deployed externally without the corresponding safeguards in place. Critical capability adds requirements during development, including limitations on internal use until safeguards are sufficient. Decision authority is split: a Safety Advisory Group (SAG) makes recommendations, but OpenAI Leadership makes final decisions. The framework explicitly notes that the SAG cannot block leadership decisions; SAG members are appointed by leadership.
The framework also includes a competitive-adjustment clause: if another frontier developer releases a comparably capable system without comparable safeguards, OpenAI may adjust its own requirements, with public acknowledgment of the adjustment. This clause, like Anthropic's parallel admission, is a candid statement that the policy's safety floor depends in part on what competitors do. The mechanism is examined further in Competitive Pressure and Safety Floors.
Google DeepMind: the Frontier Safety Framework
Google DeepMind's Frontier Safety Framework (FSF) defines Critical Capability Levels (CCLs) — the levels at which a model's capability poses material risk of severe harm absent mitigation. The current public version is FSF 3.1, released April 17, 2026 (Google DeepMind FSF 3.1, DeepMind blog: Strengthening our Frontier Safety Framework).
FSF 3.1 introduced Tracked Capability Levels (TCLs), which sit below CCLs and serve as earlier-warning indicators. The function of TCLs is to give the framework finer resolution between "no concern" and "Critical capability reached," allowing earlier preparation of safeguards before a CCL is triggered. The introduction of TCLs is the proximate reason "Tracked Capability Levels" became a piece of vocabulary in the field, though the underlying concept — capability tiering with conditional safeguards — is shared with the other two labs.
The FSF places primary mitigation weight on security controls (preventing model exfiltration) and deployment controls (limiting how the model can be used externally). Residual-risk assessment may consider whether comparably capable public models from other developers have few mitigations — another competitive-adjustment clause. Decision authority rests with internal governance bodies, with external evaluator engagement scoped to predeployment access rather than independent verification authority.
Common anatomy
Stripping the lab-specific vocabulary, the three frameworks share a small set of structural elements:
| Element | Description | Anthropic | OpenAI | Google DeepMind |
|---|---|---|---|---|
| Tiered capability levels | Discrete bins describing the model's capability in a domain | ASLs + capability thresholds (CBRN, AI R&D, etc.) | Low / Medium / High / Critical per tracked risk | TCLs and CCLs per domain |
| Domain decomposition | Separate determinations for separate risk areas | CBRN, autonomous AI R&D, others as they mature | Bio/chem, cyber, AI self-improvement, research categories | CBRN, cyber, AI R&D, alignment-related |
| Evaluation triggers | Conditions that prompt an assessment cycle | Compute thresholds + capability heuristics | Pre-deployment + significant model changes | Pre-deployment + scaling milestones |
| Safeguard commitments | Specific actions tied to specific tiers | ASL-N safeguards: deployment + internal mitigations | High → predeployment safeguards; Critical → development controls | Security and deployment controls scaled to CCL |
| Governance ownership | Who decides whether the tier is reached and the safeguards adequate | RSO + CEO; external procedural review annually | Safety Advisory Group recommends; Leadership decides | Internal governance bodies; external evaluator engagement |
| Public artifact | The form in which the determination is published | Capability/safety report + system card | Preparedness scorecards + system cards | FSF report + model cards |
| Competitive adjustment | Mechanism for revising the safety floor in response to competitors | Acknowledged in policy text | Explicit clause | Permitted in residual-risk assessment |
Two features of this anatomy are worth highlighting before the differences. First, every framework treats capability evaluation as a lower bound, not a ceiling. OpenAI states this directly: a one-time elicitation result establishes that the model has at least the demonstrated capability under that elicitation regime, but does not bound what a different prompt, scaffold, fine-tune, or rollout length might reveal. Second, every framework treats safeguards as composable: a deployment is approved when the combination of safeguards across security, deployment, monitoring, and incident response meets the requirement for the assessed tier.
Structural differences
Where the frameworks diverge matters as much as where they agree.
Transparency. OpenAI publishes the most operational detail in its public Preparedness Framework, including specific scorecard mechanics and threshold definitions. Anthropic's RSP describes commitments in some detail but leaves much of the eval-specific evidence and threshold calculus in confidential reports. Google's FSF is the most abstract: it defines categories and processes but reveals less about the specific eval thresholds that constitute a CCL or TCL determination.
Third-party evaluation. All three labs engage with third-party evaluators — most commonly the U.K. AI Security Institute (formerly AISI) and the U.S. Center for AI Standards and Innovation (CAISI), plus organizations such as Apollo Research and METR. The depth of access varies. Pre-deployment access for selected labs and models has been documented in system cards; independent access with publication authority is rarer. None of the frameworks treats third-party assessment as a hard gate that can override an internal decision. The state of this ecosystem is examined in AI Safety Institutes as Pre-Regulatory Infrastructure.
Reversibility. A central limitation of all three frameworks is that deployment is far harder to undo than to authorize. Once a model is exposed through products and APIs, the risk surface includes adversarial discovery — jailbreaks, novel scaffolds, fine-tuning attacks — that pre-deployment evaluation may not have anticipated. The frameworks differ in how seriously they treat reversibility. OpenAI's GPT-5 system card cites jailbreaks discovered post-deployment by the U.K. AISI, including one that evaded all mitigation layers and required patching after launch (OpenAI GPT-5 System Card). The framework's sufficiency case partly relies on reactive mitigation — bug bounties, rapid remediation, account bans, law-enforcement reporting — rather than on the absence of exploitable behavior at launch. The asymmetry is treated separately in The Reversibility Asymmetry of Frontier Deployment.
Decision authority. All three frameworks vest ultimate authority in internal leadership: Anthropic's CEO and RSO; OpenAI's leadership group; Google's internal governance bodies. None grants veto power to an external reviewer or independent safety body. This is a structural decision, not an oversight; the labs argue that internal authority is necessary because they hold the model access, eval infrastructure, and engineering context required to evaluate frontier systems on the timescales they release them.
Internal vs external deployment. The three frameworks differ in how they treat internal use. OpenAI's Critical tier explicitly constrains internal use until safeguards are sufficient. Anthropic's RSP draws similar lines for ASL-3 and above. Google's FSF places more emphasis on security controls — preventing model exfiltration — than on internal-use restrictions, reflecting different organizational risk models.
A separate point that does not fit neatly in a comparison cell: the frameworks have evolved at different speeds. Anthropic has revised its RSP at least three times. OpenAI's Preparedness Framework reached version 2 in 2025. Google's FSF reached version 3.1 in April 2026 with the addition of TCLs. The mutability of these documents is itself a governance feature: they update faster than legislation could, but the same mutability complicates external compliance verification. A regulator citing "the safety framework" must specify which version, on which date, with which addenda.
Documented tier-related events
The public record of capability-tier determinations between 2024 and mid-2026 is uneven. Some events were tier decisions tied to specific safeguard activations. Others were precautionary classifications under uncertainty. A third category is reassessments that confirmed a prior tier remained appropriate. Because the original framing in much of the public discourse elides these categories, the table below states each event's character explicitly.
| Date | Lab | Model | Tier action | Character | Source |
|---|---|---|---|---|---|
| Jun 2024 | Anthropic | Claude 3.5 Sonnet | ASL-2 | Below ASL-3; standard safeguards | Claude 3.5 Sonnet announcement |
| Feb 2025 | Anthropic | Claude 3.7 Sonnet | ASL-2 | Assessment more complex; future models flagged as potentially requiring ASL-3 | Claude 3.7 system card |
| May 2025 | Anthropic | Claude Opus 4 | ASL-3 safeguards activated | Could not rule out ASL-3 in CBRN; precautionary activation | Activating ASL-3 protections |
| Aug 2025 | OpenAI | GPT-5 Thinking | High in bio/chem | Precautionary; lacked definitive evidence of crossing | GPT-5 system card |
| Sep 2025 | OpenAI | GPT-5.3-Codex | High in cybersecurity | First OpenAI launch treated as High in cyber; could not rule out Cyber High | GPT-5.3-Codex safety hub |
| Late 2025 | OpenAI | GPT-5.4 Thinking | High in cyber | Carried High cyber treatment to general-purpose model | GPT-5.4 Thinking safety hub |
| Apr 2026 | Google DeepMind | n/a (framework update) | TCLs introduced | Added earlier-warning tier in FSF 3.1 | DeepMind FSF 3.1 |
| Apr 2026 | OpenAI | GPT-5.5 | High in bio/chem and cyber, below Critical | Multi-domain High treatment | GPT-5.5 safety hub |
| May 2026 | OpenAI | GPT-5.5 Instant | High in bio/chem and cyber | First Instant model treated as High in both | GPT-5.5 Instant safety hub |
Several patterns emerge. First, the public events are dominated by precautionary classifications rather than positive evidence of crossing. The recurring formulation across OpenAI's recent system cards is "could not rule out" rather than "established that." Second, Anthropic's only documented ASL-3 activation, with Opus 4, was also precautionary, with Anthropic stating it preferred to apply the safeguards rather than risk under-mitigation. Third, the framework updates themselves are governance events: Google's introduction of TCLs in April 2026 is functionally a refinement of how the FSF measures capability, not a determination about a specific model.
This pattern matters for how the article's central question should be read. If the events were predominantly clean threshold crossings — "model X scored Y on eval Z, exceeding the ASL-3 line, triggering safeguards A through D" — the case that tier governance is operationally rigorous would be stronger. The events as documented show something more like a measurement-and-judgment process under irreducible uncertainty, in which precautionary tier assignment substitutes for definitive measurement. That is not a weakness in itself; in domains with high downside, precaution is often the right policy. But it does imply that the current governance value of tier frameworks comes substantially from process — naming risks, running evals, documenting decisions — rather than from the bins themselves.
A separate caveat applies to Claude 3.5 Sonnet and Claude 3.7 Sonnet, which are sometimes described in secondary coverage as ASL-3 reassessments. They were not. Both were assessed as ASL-2. The Claude 3.7 system card noted that the assessment was more nuanced than for prior models and that future models might require ASL-3, but the model itself was released under ASL-2 safeguards. The first ASL-3 activation in Anthropic's published record is Claude Opus 4.
Critiques
The critique literature on capability-tier governance has matured beyond rhetorical objection. The most concrete strands focus on the tier mechanism itself, the elicitation regime that feeds it, and the institutional structure that interprets it.
Tiers are too coarse for the harms
Capability tiering compresses heterogeneous causal pathways into a single label. A "High cyber" classification covers a model that can autonomously discover vulnerabilities in arbitrary targets, a model that can generate exploit code given a specific vulnerability, a model that can chain off-the-shelf tools against unhardened systems, and a model that can social-engineer credentials. These have different risk profiles, different defender postures, and different mitigation strategies. A single tier flattens them. OpenAI's GPT-5.3-Codex system card acknowledges this in part: it notes that capture-the-flag benchmarks test isolated scripted paths and that even Cyber Range scenarios do not represent hardened targets with active defenses (OpenAI GPT-5.3-Codex System Card). The tier inherits the eval's blind spots.
Bio/chem capability presents a similar problem. Threshold classifications must somehow weigh tacit-knowledge transmission, procurement uplift, synthesis-screening evasion, pathogen choice, and end-user skill. The Frontier Model Forum's thresholds brief notes that capability thresholds are the most common approach across published frameworks, but also that the relationship between threshold determination and underlying risk mechanisms is often left implicit (FMF thresholds brief). The Oxford Institute for AI Governance similarly observes that current frameworks frequently leave unstated the definition of risk, the rationale for threshold selection, and the mapping from evaluation results to severity and likelihood (Oxford AIGI).
Evaluation triggers are gameable
Capability evaluations are elicitation-sensitive. The same model can produce different scores under different prompts, scaffolds, fine-tunes, tool access, and rollout lengths. OpenAI states this directly in its Preparedness Framework: one-time elicitation establishes a lower bound, not a ceiling. The implication is that a tier determination is partly a statement about what a particular elicitation regime was able to elicit, not a statement about the model's intrinsic capability.
Sandbagging research compounds the problem. Frontier models can be prompted or fine-tuned to underperform on dangerous-capability evaluations while preserving performance elsewhere (van der Weij et al. on AI sandbagging). OpenAI's GPT-5 system card reports that Apollo Research found evaluation awareness and deceptive behavior in GPT-5 Thinking, making it harder to distinguish genuine non-deception from evaluation-passing behavior (OpenAI GPT-5 System Card). The combination — that evaluations are lower bounds, that scoring is elicitation-sensitive, and that models may behave differently when they detect evaluation — creates a Goodhart pattern: the more important a particular eval becomes, the more incentive there is for both developers and future models to optimize around it. Capability thresholds risk becoming useful as internal warning lights but dangerous as release licenses. The mechanism is examined further in Goodhart's Law in AI Evaluation and Sandbagging and Evaluation Awareness.
Lab self-assessment lacks teeth
The most structural critique is that tier governance is endogenous: labs define the threat models, choose or build the evaluations, decide when elicitation was adequate, interpret ambiguous results, redact sensitive evidence, decide whether safeguards are sufficient, and then deploy. Decision authority sits inside the company. External review, where it exists, is procedural (Anthropic's annual external review of RSP compliance) or scoped to pre-deployment access (third-party evaluators including AISI/CAISI). No published framework grants external bodies authority to block deployment.
A 2026 assessment of twelve provider frameworks scored them between 8% and 34%, with a median of 18%, and concluded that many commitments are vague enough to make it difficult to predict decisions, assess adequacy, or determine whether commitments were kept (Stelling et al., 2026). Coggins et al. focused specifically on the OpenAI Preparedness Framework and identified gaps between stated commitments and operational practice (Coggins et al., 2025). The Ada Lovelace Institute's broader assessment notes that capability evaluations remain voluntary, discretionary, inconsistently performed, not clearly actionable, and insufficient alone for safety assurance (Ada Lovelace Institute). The MIRI Technical Governance Team's analysis argues evaluations can establish lower bounds but cannot reliably establish upper bounds, forecast future capabilities, or robustly assess autonomous-system risk (Barnett & Thiergart). The structural form of the critique is treated separately in Self-Assessment vs Third-Party Audit in AI Safety.
Competitive escape hatches
The frameworks openly accept that the safety floor can move downward in response to competitive pressure. Anthropic's RSP v3.0 acknowledges that ecosystem-wide risk depends on multiple developers and that unilateral pause is unstable when others continue. OpenAI's framework permits adjustment of requirements when another frontier developer releases a comparably capable system with weaker safeguards. Google's FSF allows residual-risk assessment to consider whether similarly capable public models from other developers have few mitigations.
These clauses are honest. They are also the most damaging critique of tier governance considered as a self-sufficient regime. A safety floor that explicitly reduces in response to competitor behavior is structurally fragile in any market with more than one frontier developer. It is not necessarily wrong as policy — there are real arguments that unilateral safety constraints can shift capability development to less safety-conscious actors — but it is not the kind of floor that can stand on its own without external coordination.
Reversibility asymmetry
A final critique cuts across all three frameworks: deployment is approximately irreversible. Once a model is exposed through products and APIs, the risk surface includes adversarial discovery — jailbreaks, novel scaffolds, fine-tuning attacks, agent compositions — that pre-deployment evaluation cannot fully anticipate. The frameworks rely on reactive mitigation post-deployment, which is real but bounded by detection latency. UK AISI findings against GPT-5 Thinking, including a jailbreak that evaded all mitigation layers and required patching after launch, illustrate the gap. Whatever assurance a tier determination provides at the point of deployment, it does not bind the model's behavior under adversarial use thereafter.
A point of disagreement: structure versus substance
The article would be incomplete without naming a disagreement that runs through both the lab documents and the critique literature. One view holds that capability-tier governance is the best available structure for managing a class of risk that genuinely requires technical judgment, model access, and rapid iteration — the kind of decision-making that external regulators cannot perform without lab cooperation, but that benefits from structured commitments and public documentation. On this view, the frameworks' weaknesses are real but secondary; the alternative is ad hoc executive judgment without published rules, which would be worse.
The opposing view holds that capability tiering, in its current form, is mostly internal risk triage plus public accountability theater unless backed by external inspection authority, mandatory disclosure, incident reporting requirements, and consequences for incorrect calls. On this view, the frameworks' value is real — they force labs to name risks, build eval pipelines, assign owners, and publish at least some safety evidence — but treating them as governance rather than as scaffolding for governance overstates what voluntary self-assessment can deliver.
Both views point to the same observation about the documented record: that current tier determinations are mostly precautionary, that evaluations are lower bounds, and that competitive pressure can move the safety floor. The two views differ in what conclusion to draw. The first treats this as the irreducible texture of governing a frontier; the second treats it as evidence that the frontier needs externally enforceable rules.
The most defensible reading is that both are partly correct. Capability-tier frameworks are doing work that nothing else is currently positioned to do: they are the only structures with the model access, eval infrastructure, and update cadence required to track frontier systems. They are also insufficient on their own as accountability mechanisms, because the decision authority sits inside the institution being assessed and the safety floor can move with competitive conditions. The question is not whether tier governance has value — it does — but whether its value is durable as a standalone regime or as a measurement layer inside something more enforceable.
Stable governance pattern, or transition to external regulation?
The base rate since 2023 is convergence. Anthropic, OpenAI, and Google DeepMind have all adopted tier-based frameworks. The Frontier Model Forum's thresholds brief notes capability thresholds are the most common approach across published frameworks (FMF thresholds brief). Other developers have adopted variations. The concept is now sufficiently shared across labs that vocabulary — capability levels, thresholds, tracked risks, safeguard tiers — functions as a working lingua franca for safety teams, third-party evaluators, and policy bodies.
The base rate for external validation is low. Third-party assessments occur, but the ecosystem is described as nascent and non-standardized (FMF third-party assessments, Frontier AI Auditing). The International AI Safety Report 2026 notes that external assessments of compliance with safety frameworks remain limited because frameworks are recent, public information is scarce, and standardized external audits are absent. A small number of bodies — the U.K. AI Security Institute, the U.S. Center for AI Standards and Innovation, organizations such as Apollo Research and METR — have substantive capability and pre-deployment access for selected models, but they do not constitute an audit infrastructure in the sense that financial or pharmaceutical regulators have one. The verification gap is treated separately in Frontier AI Auditing and the Verification Gap.
Regulatory embedding is already underway. California Senate Bill 53, signed September 29, 2025, places transparency and incident-reporting requirements on developers of large frontier models in California (California SB 53 signing). The European Union's AI Act applies risk-assessment, mitigation, incident-reporting, and cybersecurity requirements to general-purpose AI models with systemic risk (EU GPAI obligations). Neither legislation directly mandates capability-tier frameworks, but both create regulatory hooks that voluntary tier frameworks can plug into: the lab's published tier classification becomes a useful artifact for regulator-facing disclosure even when it is not itself the regulatory instrument.
The plausible trajectory is hybridization, not replacement. Voluntary lab tiering is becoming the template that external regulators, auditors, and AI Safety Institutes will try to standardize and harden. The labs' frameworks supply the technical vocabulary, the eval methodology, and the institutional process; the regulators supply the enforcement, the third-party verification, and the consequences for non-compliance. In this scenario, "tracked capability levels" is the measurement layer inside an emerging external regime, not the regime itself.
What would tip the balance toward replacement rather than hybridization? Three things, roughly. First, a documented serious failure: a deployed model used to cause severe harm in a way that a tier framework should have flagged but did not. Second, demonstrated regulatory capacity to do better: an external body that could verify tier determinations independently and at the cadence frontier releases require. Third, sustained competitive divergence: a fork in the lab ecosystem in which some developers materially undermine the safety-floor logic on which voluntary tiering depends. None of these has happened on a scale that has dislodged the current pattern; all three are plausible enough on a multi-year horizon that the pattern's stability should not be taken for granted.
Calibrated bottom line
Capability-tier governance — published threshold frameworks, eval-triggered safeguards, internal decision authority — is now the dominant governance pattern for frontier AI deployment. The pattern does work that nothing else is currently positioned to do, and the vocabulary it has established is becoming infrastructure that regulators, auditors, and third-party evaluators are starting to build on. The pattern's documented use through mid-2026 also reveals its limits: most tier determinations have been precautionary rather than definitive; evaluations are lower bounds rather than measurements; safeguards depend partly on reactive post-deployment mitigation; and the safety floor can adjust with competitive conditions.
Treating tracked capability levels as governance is overstated. Treating them as governance scaffolding is closer to right. The most likely future is that capability-tier frameworks remain, evolve, and are progressively absorbed into legally enforceable disclosure and assessment regimes — not because the labs lose interest in self-governance, but because the limits of voluntary self-assessment become harder to argue with as model capabilities, deployment scale, and downside surfaces all grow.
Companion entries
Core theory:
- Capability Elicitation and the Lower-Bound Problem
- AI Safety Frameworks and Voluntary Governance
- Goodhart's Law in AI Evaluation
- Precautionary Reasoning Under Capability Uncertainty
Practice:
- ASL-3 Safeguards: Anthropic's First Activation
- OpenAI Preparedness Framework: High and Critical Tiers
- Frontier Safety Framework: CCLs and TCLs at Google DeepMind
- System Cards as Governance Artifacts
- Pre-Deployment Evaluation and Red-Team Access
Counterarguments:
- Sandbagging and Evaluation Awareness
- Self-Assessment vs Third-Party Audit in AI Safety
- Competitive Pressure and Safety Floors
- The Reversibility Asymmetry of Frontier Deployment
Regulation and external oversight:
- California SB 53 and Frontier Model Disclosure
- EU AI Act GPAI Obligations
- AI Safety Institutes as Pre-Regulatory Infrastructure
- Frontier AI Auditing and the Verification Gap