The same week Anthropic refused the Pentagon’s demand to remove its AI safeguards, the company quietly rewrote the document those safeguards were built on.
Both things happened. They’re both worth understanding.
What the original promise was
In 2023, Anthropic published the first version of its Responsible Scaling Policy (RSP), a public document that laid out the conditions under which the company would and wouldn’t continue developing more powerful AI models.
The core commitment was specific: if Claude’s capabilities ever outpaced Anthropic’s ability to guarantee safety, the company would stop. Full pause on training and deploying new models until safety measures caught up. No exceptions carved out for competitive pressure or market conditions.
That was the promise. Anthropic was the only major AI lab that had made it.
What changed on February 24th
On Tuesday of last week, Anthropic published RSP version 3.0. The hard stop is gone.
In its place is a new framework called the Frontier Safety Roadmap, where Anthropic commits to being transparent about safety risks and publishing regular reports on where they stand. The binding “we will pause” commitment has been replaced with “we will tell you what we’re doing about it.”
Their explanation, given in a Time interview with chief science officer Jared Kaplan, is essentially this: unilateral commitments don’t make sense when competitors aren’t making the same ones. If Anthropic stops building while OpenAI and Google keep going, the developers with the weakest safety standards end up setting the pace for the whole industry. A responsible actor benching themselves doesn’t make the field safer overall.
That argument is not absurd. It’s actually a real problem in how safety governance works across competing organizations.
Why it still matters that the hard stop is gone
Here’s the thing: the original commitment was valuable precisely because it was unconditional. A promise you keep only when it’s easy isn’t really a promise. The RSP said “even if this costs us competitively, even if everyone else is moving faster, we stop.” That was the thing that made it meaningful.
RSP 3.0 replaces the unconditional commitment with a conditional one. Anthropic will delay development to ensure safety if they believe they have a significant lead over competitors, or if there’s strong evidence that competitors are also using strong safety measures. But if competitors are advancing with weaker safeguards… they’ll try to keep up.
That’s a very different position than “we stop when we can’t guarantee safety.” It’s closer to “we’ll do what we can, but we’re not unilaterally pausing if everyone else is still going.”
Anthropic explicitly says they separated what they’ll do alone from what requires industry-wide action. That’s an honest acknowledgment. But honest acknowledgment that a commitment was too hard to keep is not the same as the commitment itself.
The timing problem
Anthropic says the RSP revision and the Pentagon standoff are unrelated. The RSP 3.0 process started a year ago in February 2025, well before this week’s military contract dispute.
That’s probably true. But the rollout timing puts them in the same news cycle, and that matters for how people interpret the company’s positioning. On the same day their chief science officer told Time they couldn’t make unilateral commitments while competitors blazed ahead, their CEO was sitting across from the Defense Secretary being told to remove safeguards or lose the contract.
The two stories don’t have to be connected to sit awkwardly next to each other.
What this means in practice
For those of us using Claude for creative and business work, nothing about the RSP change affects how the tool functions today. The Constitutional AI design, the values built into the training, the things Anthropic actually refused the Pentagon… none of that changed.
What changed is the governing document that set limits on how fast and how far Anthropic would push into more powerful and potentially more dangerous territory. The external brake, the one that existed regardless of market conditions, is no longer in the policy.
There’s a version of this that’s defensible. The AI industry is genuinely a collective action problem, and the RSP’s own architects admit that voluntary self-regulation by a single company has real limits. The new policy is more honest about what one company can actually guarantee.
But as of this week, there is no longer any major AI lab with a binding public commitment to stop development if safety can’t keep up.
That’s worth knowing.