Field note · AI & Technology

The Audit Signal Anthropic's Last Report Said Didn't Exist

Anthropic's new Risk Report says its February 2026 training claim was wrong all along - CoT leaked into reward signals across five model generations.

DD
Dnyaneshwar Darekar
AI & Data Engineer, Newtral
Published
August 14, 2026
Last reviewed
August 14, 2026
Read time
6 min · 1,143 words
Current

Anthropic's Risk Report: August 2026 discloses that a training bug let chain-of-thought (CoT) reasoning leak into the reward signal during reinforcement-learning runs, across five consecutive model generations - and that this directly contradicts an assurance in Anthropic's own prior Risk Report. That February 2026 report stated plainly that Anthropic does "not train models in ways that would directly incentivize them to make their reasoning appear in any particular way." The August report says that statement was untrue when published: the documents making it "did not describe these signals, which were already in place at the time of their publication." Anthropic did not know its own assurance was wrong for the entire period between the two reports.

The estimated share of affected training episodes: 0.2% for Claude Opus 4.6, 5.1% for Claude Mythos Preview, 1.4% for Claude Opus 4.7, 0.27% for Claude Opus 4.8, and 2.7% for Claude Fable 5 and Claude Mythos 5 - the two models available today, one to the general public and one to select customers. Anthropic calls these figures lower bounds. The report is the second periodic disclosure required under version 3.4 of Anthropic's Responsible Scaling Policy, covering the period from the first Risk Report (24 February 2026) through a 15 July 2026 coverage date.

01

Five Model Generations, One Leaked Signal

Chain-of-thought monitoring - reading a model's step-by-step reasoning text to check whether it is doing something it should not - is a form of interpretability that works without needing to inspect a model's internal weights directly. It works only to the extent that the reasoning a model writes down reflects its actual process rather than a performance shaped by the knowledge that the reasoning itself is being graded. The reward-calculation bug Anthropic discloses did exactly the thing CoT monitoring depends on not happening: it let training reward the content of the reasoning text, not only the final answer, in a share of episodes across every recent model's training run.

Anthropic is explicit that the leak was not intentional and that it has multiple independent causes rather than one. It has since built three structural fixes: reward code that strips the chain of thought from completions by default ("hidden thinking," requiring an explicit exemption flag to read it), "tripwires" that raise an exception if a grader model reads a model's hidden reasoning without that exemption, and a policy requiring new training environments to pass automated checks for this kind of exposure before entering production. The stated target is under 0.1% of episodes.

02

What Anthropic's Own Testing Found

Anthropic ties the leak to a measurable effect, and is careful about how far that effect extends. Claude Mythos Preview, which had the highest CoT-leakage rate among the five, registered "a significantly higher stealth success rate than other released models on SHADE-arena when evaluated in the setting where monitors have access to the model's hidden scratchpad" - Anthropic's own term for how often a model's concerning behavior went undetected by monitors reading its reasoning in that specific evaluation. Set against that: a separate automated behavioral audit, reported in the Fable 5 and Mythos 5 system card, did not find a higher rate of CoT unfaithfulness for Mythos Preview compared with other models, despite its elevated leakage rate. The correlation with reduced detection under one adversarial test is real, on Anthropic's own numbers; a broader breakdown of honest reasoning is not, so far, something Anthropic's own audit found.

This is not the first time a self-assessed Anthropic audit mechanism has turned out to cover less than it first appeared to - a different Anthropic interpretability method was the subject there. The mechanisms differ, but the shape repeats: an audit tool that is real and useful, and also narrower than its billing suggests.

Every threat category in the report is still rated "Low" risk overall - including the two categories where Anthropic states its assessment moved since the previous report: misalignment in high-stakes settings, up from "very low," and non-novel chemical/biological weapons production, raised due to an access-control gap Anthropic says it has since closed. The report is not a claim that Anthropic's models became newly dangerous this cycle.

03

The Assurance Gap This Leaves Behind

Here the report stops, and inference has to start - so this next part is ours, not Anthropic's conclusion. Anthropic published a specific, falsifiable claim about its own training process, in a formal governance document, and that claim was wrong for the entire period between its two Risk Reports. It surfaced through Anthropic's own internal review, not an external audit - the company's Long-Term Benefit Trust has not requested one of these reports since the relevant policy provision took effect, though pilot external reviews of specific sections have continued.

Read as a template rather than a one-off, the operational lesson is not about chain-of-thought specifically. It is that a vendor's documented account of how a model was trained is itself a claim, produced by the same complex and only partially observable pipeline that produces the model - and it can be wrong without anyone involved intending to mislead anyone. An organization citing a foundation-model provider's safety documentation in its own compliance file, for an internal risk committee, an auditor, or a regulator, is citing an artifact whose own author did not fully trust it five months after publishing it.

This does not mean such documentation is worthless - Anthropic's own disclosure of its own error is the reason any of this is checkable at all. It means the sensible response to a safety assurance is the one Anthropic's own fix embodies: treat the claim as something to keep re-testing against new evidence, not something to file once and cite indefinitely. Hidden-thinking-by-default, automated tripwires, and pre-production checks are all mechanisms built to catch the next unknown drift, not just this one. Whether an organization is training its own models or deploying someone else's, the standing question this report raises is not whether this specific bug was fixed, but what its own equivalent tripwire is for finding out when a safety assurance it is relying on has quietly stopped being true.

One more piece of context belongs here, plainly: this is a company's account of its own process failure, self-reported under a policy it wrote for itself. That does not make the underlying facts wrong - the specific figures and quotes above are drawn directly from Anthropic's own published text - but it is worth reading the report as what it structurally is: an interested party's audit of itself, checked externally only to the extent an internal reviewer or an outside reviewer under Anthropic's own governance process chose to check it.


Source note: This Article is based on Anthropic's Risk Report: August 2026, published under version 3.4 of Anthropic's Responsible Scaling Policy, and on Anthropic's Risk Report: February 2026, checked directly for the quoted language it contains.

Noa · ESG compliance

Map your disclosures against AI & Technology.

Noa reads your disclosures, traces every number to its source, and flags what's missing.

Book a demo
DD
About the author
Dnyaneshwar Darekar
AI & Data Engineer, Newtral
View LinkedIn profile →