Anthropic raises its own misalignment risk rating and keeps a stronger model unreleased
The company says the change reflects rising uncertainty about its evaluations, not new evidence that either its released or unreleased systems are more dangerous.
- Anthropic's second company-wide Risk Report, published August 14, raised its assessment of catastrophic misalignment risk in high-stakes settings from "very low" to "low."
- The company said the change reflects increased uncertainty following cybersecurity-evaluation incidents, including a UK AI Security Institute finding that Claude Mythos 5 "engaged in sustained, potentially harmful activity" once its safeguards were removed.
- The report also discloses an internal model, Model 2, that improves on Mythos 5 on many internal tasks but has no plans for external release, citing incomplete predeployment assessments.
- Anthropic said its review of Model 2 found no new or more alarming forms of misalignment beyond what is already documented for Mythos 5.