Skip to main content
WireRead
Back to all news

Anthropic

The Fable 5 jailbreak wasn't unique — every frontier model shared it

The export order that pulled Claude Fable 5 traced to one cyber jailbreak. Anthropic's later investigation found the same bypass worked across the frontier.

By , Editor-in-Chief · WireReadVerified July 2026

The answer

Anthropic's later investigation found the Fable 5 cyber jailbreak worked on rival frontier models too, not just Claude.

The US export-control order that took Claude Fable 5 and Mythos 5 offline on 12 June began with a narrow, specific finding. Researchers at Amazon had discovered a way to bypass Fable 5's safeguards and coax the model into identifying software vulnerabilities — in at least one case producing working demonstration exploit code. That single result was enough to trigger roughly ninety minutes' notice and a fortnight of darkness for two of the most capable models on the market. What the intervening weeks established, however, is that the flaw was never Anthropic's alone.

What actually triggered the export order

The immediate cause was cyber capability, not a general safety failure. Amazon's researchers showed that Fable 5 could be prompted past its guardrails to do vulnerability research — the dual-use capability that sits at the centre of the administration's frontier-model concerns. Anthropic was given minutes, not days, to comply; Fable 5 and Mythos 5 went dark on 12 June and stayed offline for more than two weeks while the company and the government negotiated the terms of their return. The episode read less like a routine safety review than a live test of how far Washington's leverage over a frontier lab actually reaches. Two weeks is a long time for two of the highest-traffic models on the market to sit dark, and every additional day raised the commercial cost of getting the fix — and the politics — right.

Anthropic's AI models are back online after a two-week government standoff—settling the company and administration into a fragile truce.

Source: Fortune · 1 July 2026

Anthropic's investigation: the flaw is industry-wide

During the standoff, a group of cybersecurity experts wrote an open letter to the administration making a simple argument: the vulnerability was not unique to Anthropic, and other leading models had the same weakness. Anthropic's own follow-up work, reported in early-July coverage, put names to that claim. The company said its subsequent investigation found the same bypass worked on Claude Haiku 4.5, Sonnet 4.6, Opus 4.6, Opus 4.7 and Opus 4.8, as well as OpenAI's GPT-5.4 and GPT-5.5 and Moonshot's Kimi K2.7 — an entire class of jailbreak spanning the frontier, not a defect specific to Claude. Three labs, eight models, one shared root cause: a very different picture from the one the export order implied when it singled out Fable 5 alone.

  • Claude: Haiku 4.5, Sonnet 4.6, Opus 4.6, Opus 4.7 and Opus 4.8
  • OpenAI: GPT-5.4 and GPT-5.5
  • Moonshot: Kimi K2.7

Gradually, various agency heads approved of the changes, and on July 1 the models were released, the official said.

Source: Axios · 3 July 2026

Why singling out one lab is the wrong fix

If the jailbreak is industry-wide, switching off one company's models is closer to arbitrary than surgical. The capability that worried Amazon's researchers exists across the frontier; taking Fable 5 out of the market did nothing about GPT-5.5 or Kimi K2.7. To bring its models back, Anthropic trained a new safety classifier specifically to block the technique, published a draft 'AI jailbreak severity framework' to grade such findings, and opened a HackerOne programme inviting researchers to report cyber jailbreaks. Those are shared-defence moves — the kind of standard that only works if every lab adopts it.

That is the strategic point beneath the episode. The durable fix for a class of vulnerability is a common defensive standard and a disclosure pipeline, not the selective de-platforming of whichever lab a regulator happened to notice first. The same dynamic is now pushing Washington towards voluntary pre-release testing standards, and it is sharpening a familiar objection: heavy-handed reach into US labs does nothing to constrain the cheaper open-weight models that carry the same flaws and answer to no export order.

The Trump administration's latest restrictions on private AI model releases are ramping up the push for open-source alternatives.

Source: The Hill · 30 June 2026

Frequently asked questions

Was the Fable 5 jailbreak unique to Claude?
No. Cybersecurity experts argued during the standoff that the flaw was shared, and Anthropic's subsequent investigation, reported in early July, found the same bypass worked across the frontier — including on OpenAI's GPT-5.5 and Moonshot's Kimi K2.7.
Which models did Anthropic's investigation say were affected?
Per early-July reporting, Anthropic found the bypass worked on Claude Haiku 4.5, Sonnet 4.6, Opus 4.6, Opus 4.7 and Opus 4.8, plus GPT-5.4, GPT-5.5 and Kimi K2.7 — a class of jailbreak rather than an Anthropic-specific defect.
Why did the US pull Claude Fable 5 in the first place?
Amazon researchers found a way to bypass Fable 5's safeguards to identify software vulnerabilities, in one case producing demonstration exploit code. That cyber finding triggered the 12 June export-control order and roughly two weeks offline.
What did Anthropic do to bring the models back?
It trained a new safety classifier to block the jailbreak technique, published a draft 'AI jailbreak severity framework', and opened a HackerOne programme for researchers to report cyber jailbreaks — the conditions for restoring access.
Does this affect open-source models too?
Reporting suggests so. Open-source advocates warn that clamping down on US labs does nothing about cheaper open-weight models that share the same vulnerabilities and sit outside any export order.

Sources

← All news