Monitoring Monitorability
Melody Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Wei, Marcus Williams, Benjamin Arnav, Joost Huizinga, Ian Kivlichan, Amelia Glaese, Jakub Pachocki, Bowen Baker
Abstract
Observability into the decision making of modern AI systems may be required to safely deploy increasingly capable agents. Monitoring the chain-of-thought (CoT) of today's reasoning models has proven effective for detecting misbehavior. However, this "monitorability" may be fragile under different training procedures, data sources, or even continued system scaling. To measure and track monitorability, we propose three evaluation archetypes (intervention, process, and outcome-property) and a new monitorability metric, and introduce a broad evaluation suite. We demonstrate that these evaluations can catch simple model organisms trained to have obfuscated CoTs, and that CoT monitoring is more effective than action-only monitoring in practical settings. We compare the monitorability of various frontier models and find that most models are fairly, but not perfectly, monitorable. We also evaluate how monitorability scales with inference-time compute, reinforcement learning optimization, and pre-training model size. We find that longer CoTs are generally more monitorable and that RL optimization does not materially decrease monitorability even at the current frontier scale. Notably, we find that for a model at a low reasoning effort, we could instead deploy a smaller model at a higher reasoning effort (thereby matching capabilities) and obtain a higher monitorability, albeit at a higher overall inference compute cost. We further investigate agent-monitor scaling trends and find that scaling a weak monitor's test-time compute when monitoring a strong agent increases monitorability. Giving the weak monitor access to CoT not only improves monitorability, but it steepens the monitor's test-time compute to monitorability scaling trend. Finally, we show we can improve monitorability by asking models follow-up questions and giving their follow-up CoT to the monitor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a5e64f4-c6a3-4c8b-9dc7-c992873b804fCited by top-tier papers6
- How does information access affect LLM monitors' ability to detect sabotage?Rauno Arike, Raja Moreno, Rohan Subramani, Shubhorup Biswas et al.ICML 2026 · 11 citations
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use AgentsJaylen Jones, Zhehao Zhang, Yuting Ning, Eric Fosler-Lussier et al.ICML 2026 · 9 citations
- Outcome Rewards Do Not Guarantee Verifiable or Causally Important ReasoningQinan Yu, Alexa Tartaglini, Peter Hase, Carlos Guestrin et al.ICML 2026 · 4 citations
- Monitorability as a Free Gift: How RLVR Spontaneously Aligns ReasoningZidi Xiong, Shan Chen, Himabindu LakkarajuICML 2026 · 3 citations
- Hallucination Detection from Structural Reasoning ModelJianbo Sun, Pengkun YangICML 2026
Builds on14
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan et al.ICML 2026 · 175 citations
- SCRUPLES: A Corpus of Community Ethical Judgments on 32, 000 Real-Life AnecdotesNicholas Lourie, Ronan Le Bras, Yejin ChoiAAAI 2021 · 147 citations
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 137 citations
Related papers
- Reasoning Models Struggle to Control their Chains of ThoughtChen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He et al.ICML 2026 · 11 citations
- Output Supervision Can Obfuscate the Chain of ThoughtJacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud et al.ICLR 2026 · 10 citations
- Reasoning Models Sometimes Output Illegible Chains of ThoughtArun JoseNeurIPS 2025 · 11 citations
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky et al.NeurIPS 2025 · 50 citations
- All Code, No Thought: Language Models Struggle to Reason in Ciphered LanguageShiyuan Guo, Henry Sleight, Fabien RogerICLR 2026 · 5 citations
