CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
Benjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky, Hannes Whittingham, Mary Phuong
Abstract
As AI models are deployed with increasing autonomy, it is important to ensure they do not take harmful actions unnoticed. As a potential mitigation, we investigate Chain-of-Thought (CoT) monitoring, wherein a weaker trusted monitor model continuously oversees the intermediate reasoning steps of a more powerful but untrusted model. We compare CoT monitoring to action-only monitoring, where only final outputs are reviewed, in a red-teaming setup where the untrusted model is instructed to pursue harmful side tasks while completing a coding problem. We find that while CoT monitoring is more effective than overseeing only model outputs in scenarios where action-only monitoring fails to reliably identify sabotage, reasoning traces can contain misleading rationalizations that deceive the CoT monitors, reducing performance in obvious sabotage cases. To address this, we introduce a hybrid protocol that independently scores model reasoning and actions, and combines them using a weighted average. Our hybrid monitor consistently outperforms both CoT and action-only monitors across all tested models and tasks, with detection rates twice higher than action-only monitoring for subtle deception scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6860eb7a-6dbf-4444-baca-da3149d075e5Cited by top-tier papers10
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-ThoughtSiddharth Boppana, Annabel Ma, Max Loeffler, Raphaël Sarfati et al.ICML 2026 · 30 citations
- Reliable Weak-to-Strong Monitoring of LLM AgentsNeil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich et al.ICLR 2026 · 28 citations
- Monitoring MonitorabilityMelody Guan, Miles Wang, Micah Carroll, Zehao Dou et al.ICML 2026 · 26 citations
- Early Signs of Steganographic Capabilities in Frontier LLMsArtur Zolkowski, Kei Nishimura-Gasparian, Robert McCarthy, Roland S. Zimmermann et al.ICLR 2026 · 22 citations
- How does information access affect LLM monitors' ability to detect sabotage?Rauno Arike, Raja Moreno, Rohan Subramani, Shubhorup Biswas et al.ICML 2026 · 11 citations
Builds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan et al.ICML 2026 · 175 citations
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 137 citations
Related papers
- Output Supervision Can Obfuscate the Chain of ThoughtJacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud et al.ICLR 2026 · 10 citations
- All Code, No Thought: Language Models Struggle to Reason in Ciphered LanguageShiyuan Guo, Henry Sleight, Fabien RogerICLR 2026 · 5 citations
- Is it Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning EffortXinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell et al.ICLR 2026 · 27 citations
- Large language models can learn and generalize steganographic chain-of-thought under process supervisionRobert MC Carthy, Joey Skaf, Luis Ibañez-Lissen, Vasil Georgiev et al.NeurIPS 2025 · 29 citations
- Reasoning Models Struggle to Control their Chains of ThoughtChen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He et al.ICML 2026 · 11 citations
