Monitoring Monitorability
Melody Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Wei, Marcus Williams, Benjamin Arnav, Joost Huizinga, Ian Kivlichan, Amelia Glaese, Jakub Pachocki, Bowen Baker
摘要
Observability into the decision making of modern AI systems may be required to safely deploy increasingly capable agents. Monitoring the chain-of-thought (CoT) of today's reasoning models has proven effective for detecting misbehavior. However, this "monitorability" may be fragile under different training procedures, data sources, or even continued system scaling. To measure and track monitorability, we propose three evaluation archetypes (intervention, process, and outcome-property) and a new monitorability metric, and introduce a broad evaluation suite. We demonstrate that these evaluations can catch simple model organisms trained to have obfuscated CoTs, and that CoT monitoring is more effective than action-only monitoring in practical settings. We compare the monitorability of various frontier models and find that most models are fairly, but not perfectly, monitorable. We also evaluate how monitorability scales with inference-time compute, reinforcement learning optimization, and pre-training model size. We find that longer CoTs are generally more monitorable and that RL optimization does not materially decrease monitorability even at the current frontier scale. Notably, we find that for a model at a low reasoning effort, we could instead deploy a smaller model at a higher reasoning effort (thereby matching capabilities) and obtain a higher monitorability, albeit at a higher overall inference compute cost. We further investigate agent-monitor scaling trends and find that scaling a weak monitor's test-time compute when monitoring a strong agent increases monitorability. Giving the weak monitor access to CoT not only improves monitorability, but it steepens the monitor's test-time compute to monitorability scaling trend. Finally, we show we can improve monitorability by asking models follow-up questions and giving their follow-up CoT to the monitor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- How does information access affect LLM monitors' ability to detect sabotage?Rauno Arike, Raja Moreno, Rohan Subramani, Shubhorup Biswas 等ICML 2026 · 被引用 11 次
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use AgentsJaylen Jones, Zhehao Zhang, Yuting Ning, Eric Fosler-Lussier 等ICML 2026 · 被引用 9 次
- Outcome Rewards Do Not Guarantee Verifiable or Causally Important ReasoningQinan Yu, Alexa Tartaglini, Peter Hase, Carlos Guestrin 等ICML 2026 · 被引用 4 次
- Monitorability as a Free Gift: How RLVR Spontaneously Aligns ReasoningZidi Xiong, Shan Chen, Himabindu LakkarajuICML 2026 · 被引用 3 次
- Hallucination Detection from Structural Reasoning ModelJianbo Sun, Pengkun YangICML 2026
它引用的顶会 Paper14
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan 等ICML 2026 · 被引用 175 次
- SCRUPLES: A Corpus of Community Ethical Judgments on 32, 000 Real-Life AnecdotesNicholas Lourie, Ronan Le Bras, Yejin ChoiAAAI 2021 · 被引用 147 次
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 被引用 137 次
相关 Paper
- Reasoning Models Struggle to Control their Chains of ThoughtChen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He 等ICML 2026 · 被引用 11 次
- Output Supervision Can Obfuscate the Chain of ThoughtJacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud 等ICLR 2026 · 被引用 10 次
- Reasoning Models Sometimes Output Illegible Chains of ThoughtArun JoseNeurIPS 2025 · 被引用 11 次
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky 等NeurIPS 2025 · 被引用 50 次
- All Code, No Thought: Language Models Struggle to Reason in Ciphered LanguageShiyuan Guo, Henry Sleight, Fabien RogerICLR 2026 · 被引用 5 次
