41 AI researchers call chain-of-thought monitoring fragile
TL;DR
- A 41-author paper argues that reasoning models' human-language chains of thought are a monitorable safety signal that development choices could erode.
- The authors concede CoT monitoring is imperfect and misses some misbehavior, but urge investment in it alongside existing oversight methods.
- Signatories include Yoshua Bengio, Shane Legg, Geoffrey Irving, Neel Nanda, Ethan Perez, Evan Hubinger, Dan Hendrycks and Wojciech Zaremba, with Tomek Korbak leading.
Forty-one AI safety researchers have put their names to a short paper arguing that today's reasoning models offer a peculiar and possibly fleeting safety advantage: you can read what they are thinking.
The paper, titled 'Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety' and posted to arxiv in July 2025, opens with the claim that 'AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave.' Lead author Tomek Korbak and co-signers including Yoshua Bengio, Shane Legg, Geoffrey Irving, Neel Nanda, Ethan Perez, Evan Hubinger, Dan Hendrycks and Wojciech Zaremba concede immediately that 'CoT monitoring is imperfect and allows some misbehavior to go unnoticed.'
Two of the researchers we track shared the link.
Their pitch is not that reading a model's chain of thought is a silver bullet but that the window to do it may close. 'Because CoT monitorability may be fragile,' the abstract concludes, 'we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.' A second version of the paper went up on arxiv on December 7, 2025.
Shared on Bluesky by 2 AI experts
-
Grace @gracekind.net
Originally reported by arxiv.org
Read the original article →Original headline: Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety