Anthropic's second Risk Report raises misalignment threat level to "low," discloses shelved internal model
Anthropic published its second company-wide Risk Report on August 14, 2026, raising its assessment of the risk of catastrophic harm from AI misalignment in high-stakes settings to "low" — up from "very low" in its first report — and disclosing that an internal model called Model 2, more capable than its public flagship Claude Mythos 5, has no current plans for external release.
What's new
The report, published under Responsible Scaling Policy version 3.4, breaks its findings into per-threat-model tables. The headline change is in the misalignment-in-high-stakes-settings table, where Anthropic writes its overall risk assessment is:
"Low (an increase from our previous assessment of 'very low,' in light of general increased uncertainty around recent incident disclosures related to model behavior in cybersecurity evaluations)."
Other findings from the report:
- Automated R&D risk: rated "low," but Anthropic says it has less confidence in that call than in prior reports, because its concrete task-based evaluations have "saturated" — meaning they no longer register capability gains — right as the company reports "early signs of acceleration" in AI-assisted research and development.
- Non-novel chemical/biological weapons uplift: rated "low, but higher than our previous estimate," tied to a gap in access controls for models without blocking classifiers. Anthropic says it has since remediated the gap and found no evidence of misuse in its review.
- Novel chemical/biological weapons uplift: rated "low risk, but with substantial uncertainty."
- Model 2: the report states that Claude Mythos 5 and an unreleased internal model called "Model 2" are both used heavily inside Anthropic for coding, research, and other agentic work, and that Model 2 is somewhat more capable than Mythos 5. Anthropic says it has no current plans to release Model 2 externally.
Context
Anthropic's first Risk Report was published in February 2026 alongside RSP version 3.0, when the misalignment-in-high-stakes-settings threat model was rated "very low." This second report is the first update to that rating, and the first time Anthropic has publicly acknowledged holding back an internal model that outperforms what it ships to customers.
Why it matters
The combination of a worse risk label and a saturated internal benchmark is an uncomfortable pairing: Anthropic's own tool for detecting dangerous capability gains is losing its ability to register gains at the same moment the company says capability growth may be accelerating. That's a measurement problem as much as a safety one, and the report treats it as such rather than downplaying it.
The Model 2 disclosure is separately notable. It's a rare instance of a frontier lab stating on the record that its internal capability frontier sits ahead of what it has shipped publicly, and that the gap is a deliberate choice rather than a development lag. For a company that has built its public identity around measured, safety-first releases, choosing to disclose that gap — rather than staying silent on it — is itself a signal about how Anthropic wants to be read on this question.
Corroborating sources
- Anthropic
https://www.anthropic.com/aug-2026-risk-report
“Low (an increase from our previous assessment of "very low," in light of general increased uncertainty around recent incident disclosures related to model behavior in cybersecurity evaluations).”