
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026. Essentially, they are the same underlying model with different safeguards.Fable is generally available to the users however, like its predecessor, Mythos is accessible only to vetted cybersecurity and life science organisations. Claude Fable 5.1 and Claude Mythos 5.1 were tested for chemical and biological risk apart from cybersafety, alignment, prompt injection, agentic safety and reward hacking.
\Anthropic also cleared that external testers found no critical-severity jailbreak, while acknowledging that its automated audits are less reliable especially in long-context and multi-agent settings. The release includes new transparency measures under the EU AI Act, including watermarking for AI text and a detection API for certain firms.
How Does EU AI Act Watermarking Actually Restrict AI Misuse?
Anthropic signed the EU’s Code of Practice on Transparency of AI-generated content in July 2026 along with 190 other organisations. As mentioned, models released after August 2, would carry an invisible watermark designed to hint at the chances that Claude was used for content generation. It works by influencing word selection where several plausible options exist, and is built to survive copying, pasting and light editing. The watermark does not affect the precision or content of Claude’s responses. It also does not contain information about the users, companies, or their discussions. Anthropic has launched a detection API in private review, allowing eligible groups such as regulators, EU civil society groups, media organizations, and researchers, to verify whether the text contains the watermark.
That creates a distinct boundary between accountability and prevention. The watermark helps identify AI-generated content, but does not mention that it prevents malicious use of the model. Its primary role is to provide a framework for identifying likely AI involvement and supporting compliance with the EU regulations. The question arises: do such measures constrain misuse or create a traceability layer around increasingly capable AI systems. Anthropic’s approach puts more focus on identifying AI-generated material, while separate safeguards handle harmful requests and misuse directly.
Anthropic has bolstered enterprise guardrails through its Enterprise Frontier Safeguard System, which is created to identify and respond to misuse while allowing customers to retain data on their own cloud infrastructure under zero data retention arrangements.
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1.
— Claude (@claudeai) September 1, 2026
They're the world’s most advanced models for coding and knowledge work. pic.twitter.com/8P9PSrWPi3
What Anthropic’s Safety Report Tells Us About Reward Hacking
Claude Mythos 5.1’s security testing depicted improvements across multiple alignment measures. The company said that the model was less likely with Mythos 5 to avail resources outside its testing environment, use better reasoning, or ignore explicit restrictions when pursuing a user’s objective. It also attempted and successfully carried out reward hacking at lower rates than its prior models. However, lower reward hacking does not guarantee that unwanted behavior has been removed. Anthropic said its testing still found cases where the model could surpass approvals in automatic mode classifiers.
You may like to read– OpenAI Investigation Reveals How AI Agents Breached Hugging Face Systems
The company also admitted that its automated behavioral audit has less viability in lengthy contexts and multi-agent environments. The distinction is crucial because reward hacking refers to a situation where an AI system finds ways to fulfill an evaluation without genuinely achieving what the evaluator wanted. A reduction in such behavior is optimistic but does not ensure that the model will behave safely in every environment. Anthropic judged Mythos 5.1 to have CB-1-level chemical and biological capabilities, meaning that it could intentionally help someone with a basic technical background synthesize a known weapon, while falling short of the CB-2 threshold, for replacing rare expert talent. Its cyber capabilities are the strongest among the organization’s released model but remain in the lower category of risk under its Frontier Compliance Framework.
The company is deploying Mythos 5.1 with guardrails similar to those used for Mythos 5, including constraints around research biological capabilities. Its testing also found that the model remained robust against prompt injection and unauthorized agentic requests. The release emphasizes the difficult balance facing frontier AI developers. More capable models can perform longer and more complex tasks, but there is no single benchmark which measures whether that is safe. Anthropic’s own findings show improvement in alignment while identifying gaps in testing coverage and instances where guardrails can be passed.
For users and regulators, the important highlight is how these models perform outside controlled assessments. Reduced reward hacking, stronger guardrails, and EU AI Act watermarking add the necessary protection, but none of them establish that an autonomous AI system is safe across every real-world situation.
Also Read: Claude Marshmallow and Melon Leaks Point to Opus 5.1







