Anthropic details new Fable 5 cyber-safeguards and a cross-industry AI jailbreak severity framework
Anthropic published new detail on July 2 about the cybersecurity safeguards protecting Claude Fable 5, alongside an early draft of an AI jailbreak severity framework it has built jointly with other frontier labs.
What's new
The post lays out the layered defenses Anthropic has deployed around Fable 5's cyber capabilities. At the center are safety classifiers — in Anthropic's words, "the AI systems that accompany the model that detect and block dangerous (or potentially dangerous) cybersecurity uses." Alongside the classifiers, Anthropic says it relies on access controls, model safety training, and offline monitoring as additional layers.
Anthropic also announced a HackerOne program: "We've also launched a HackerOne program where security researchers can submit potential cyber jailbreaks they discover in Fable 5 for our review." That gives outside researchers a formal channel to report new bypass techniques rather than disclosing them ad hoc.
The centerpiece of the post is a proposed industry framework for scoring the severity of AI jailbreaks. Anthropic writes that it is laying out "an early draft version of our proposed AI jailbreak severity framework, on which we've been working with our Glasswing partners" — the Project Glasswing coalition that also includes Amazon, Microsoft, and Google. The framework scores a jailbreak along four axes: capability gain (or "uplift") — how far beyond existing tools a technique takes an attacker; breadth of capability gain ("universality") — how many distinct offensive tasks the technique works on; ease of weaponization; and discoverability. Scores run from CJS-0 (Informational) up to CJS-4 (Critical), giving labs and researchers a shared vocabulary for describing how dangerous a given jailbreak actually is.
Context
Fable 5 and its sibling model Mythos 5 spent roughly two and a half weeks off-limits to much of the world. On June 12, the U.S. government applied export controls to both models, which Anthropic said required restricting access for foreign nationals inside and outside the United States. Multiple outlets reported the controls followed a jailbreak of Fable 5's cybersecurity guardrails discovered by outside researchers, which Anthropic used to help identify and patch underlying vulnerabilities. Anthropic lifted the restriction and redeployed Fable 5 globally on July 1, alongside Mythos 5 for a subset of U.S. organizations in the Glasswing program.
Why it matters
This post is Anthropic's attempt to turn a messy, headline-grabbing incident into a repeatable process. A named severity scale for jailbreaks — with input from three of its biggest rivals — signals an attempt to standardize how the industry talks about model security failures, the same way CVSS did for software vulnerabilities decades ago. If CJS scoring catches on beyond Anthropic's own models, it would give enterprises and regulators a common yardstick for comparing safety incidents across labs, rather than relying on each vendor's own framing after the fact. It also reads as a hedge against a repeat: having a public bug-bounty channel and a graded severity system gives Anthropic (and the government) a faster, more legible way to judge the next cyber jailbreak, rather than defaulting to a blanket export-control shutdown.
Corroborating sources
- Anthropic
https://www.anthropic.com/news/fable-safeguards-jailbreak-framework
“We lay out an early draft version of our proposed AI jailbreak severity framework, on which we've been working with our Glasswing partners.”
- Cnbc
https://www.cnbc.com/2026/06/30/anthropic-says-trump-admin-has-lifted-export-controls-on-claude-fable-5-and-mythos-5.html