AI Governance

Voluntary AI safety keeps changing. Operating evidence should complement it.

Anthropic rewrote its Responsible Scaling Policy rather than dropping it. The deeper issue: self-policed AI safety commitments can change anytime.

8 min read
Joe Braidwood
Joe Braidwood
Co-founder & CEO
February 2026 · 8 min read

The signal event: Anthropic did not scrap its Responsible Scaling Policy outright. It rewrote it into a more flexible framework. That matters because even the most prominent voluntary AI safety regime can still be adjusted internally, on the company’s timetable, without any independent enforcement mechanism.

Why voluntary safety keeps moving

Anthropic’s RSP remains one of the clearest public examples of frontier-model self-governance: published thresholds, escalation language, and periodic updates. But those updates also illustrate the core limitation of voluntary commitments. The framework evolves when the company decides it should evolve.

The issue is not whether Anthropic kept the acronym alive. It is that the safeguards remain self-policed and internally adjustable.

That is the broader pattern worth paying attention to. Voluntary AI commitments are authored by the same organizations whose products, timelines, and commercial incentives they are meant to constrain. They can still be thoughtful and useful. They are just not the same thing as independently checkable operating records.

When a framework can be updated, narrowed, or reinterpreted inside the same organization, buyers and regulators benefit from separately governed records of what configured controls reported in production, paired with trusted execution, testing, and coverage evidence.

The failure is architectural, not moral

The instinct is to blame the companies. To say they lacked courage, or conviction, or integrity. But that misses the structural point entirely.

Many voluntary AI safety frameworks share a structural limitation: the organization making the promise also has substantial control over interpreting whether it kept the promise. Independent review and checkable operational records can reduce that reliance, although neither alone establishes system safety.

“A promise states intent. A signed operating record can make a covered claim checkable. Neither one, by itself, proves that a system is safe.”

This isn’t unique to AI. Financial audits, clinical-trial oversight, and building inspection all separate at least some review from the party making the claim. The comparison is imperfect, but the design lesson is useful: independent review can reduce conflicts and make evidence easier to challenge.

That is why regulated adopters should treat voluntary frameworks as inputs into diligence, not as the end of diligence.

What should complement voluntary safety: operational evidence

Voluntary commitments can change under commercial and technical pressure. They should be complemented by a different mechanism: signed, scoped records of what configured controls reported, plus the testing and coverage evidence needed to assess whether the system operated within its intended boundaries.

That mechanism is verifiable evidence.

The distinction matters:

Voluntary Promises

  • Self-reported compliance status
  • Policies describing intended behavior
  • Reputation as the enforcement mechanism
  • Revocable at the company’s discretion

Verifiable Evidence

  • Signed records of reported control outcomes
  • Signer and witness provenance disclosed when present
  • Covered integrity checks without vendor-system access
  • In-scope events, declared coverage, signed attribution

Verifiable evidence means that when an organization claims its AI system ran a safety check, a signed, timestamped record identifies the configured control, covered event, reported outcome, signer, and exclusions. A third party can check the record’s format and integrity without vendor-system access. Trusted collection, routing, testing, and coverage evidence are still needed to establish actual execution and effectiveness, and each regulator, auditor, or court decides what weight to give the record.

This evidence layer complements voluntary safety. It reduces reliance on an uncorroborated promise by making bounded operating claims checkable; it does not remove the need for governance, independent assessment, or human judgment.

Operational evidence can support review; it does not create a safe harbor

Colorado is a useful case study. Its original 2024 AI Act offered a safe harbor tied to recognized risk frameworks, but that law was repealed and replaced before taking effect by SB 26-189 (“Automated Decision-Making Technology”), signed May 14, 2026. The replacement dropped that safe harbor and instead establishes a narrower transparency regime for covered automated decision-making technology (ADMT), including pre-use notice, post-adverse-outcome disclosure, and rights to correction and human review, with substantive obligations commencing January 1, 2027. Signed operating records may support a review of whether a configured control reported a decision; they do not create a legal defense or determine compliance.

The EU AI Act’s Article 12 requires automatic recording of events for high-risk AI systems to support traceability and post-market monitoring. ISO 42001 is a management-system standard, which means it focuses on evidence that governance processes are operating rather than on marketing claims. NIST AI RMF likewise centers mapping, measuring, managing, and governing risk over time.

The regulatory direction is clear enough: self-attested promises are weaker than records that can be reviewed later.

The timing matters: Colorado’s replacement law, SB 26-189, carries substantive compliance obligations from January 1, 2027. The EU AI Omnibus entered into force on July 27, 2026; relevant high-risk obligations apply from December 2, 2027 for Annex III systems and August 2, 2028 for Annex I product-embedded systems. The closer these dates get, the less persuasive unsupported safety promises become.

The verifiable era has already begun

Revisions to company-authored safety frameworks do not mean AI safety work is pointless. They do show why governance that depends entirely on vendor promises is unstable.

A complementary evidence layer is already taking shape: signed records of reported safety-control outcomes, disclosed signer and witness provenance, and records that connect covered outputs to configured controls. Framework mappings can relate those fields to ISO 42001, NIST AI RMF, the EU AI Act, or Colorado’s disclosure regime; the mapping does not establish compliance.

Organizations deploying AI in healthcare, financial services, and other high-stakes environments need governance they can inspect. Signed records can preserve what configured guardrails reported, while deployment telemetry, testing, and coverage evidence establish whether those guardrails mediated the relevant actions and worked effectively.

The voluntary era produced useful ideas and important vocabulary. But it also produced a governance style that often stops at intention rather than independently reviewable execution.

The stronger posture is to pair voluntary commitments with records, reviews, and operational evidence that survive policy revisions and marketing cycles.

Primary sources

Pango waving

Where your AI already acts

Map what is supposed to happen, where controls block, escalate, or allow an action, and what operational evidence remains afterward.

Talk to us