Reliable Weak-to-Strong Monitoring of LLM Agents
Neil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich, Paula Rodriguez, Christina Q. Knight, Zifan Wang
Abstract
We stress test monitoring systems for detecting covert misbehavior in autonomous LLM agents (e.g., secretly sharing private information). To this end, we systematize a monitor red teaming (MRT) workflow that incorporates: (1) varying levels of agent and monitor situational awareness; (2) distinct adversarial strategies to evade the monitor, such as prompt injection; and (3) two datasets and environments -SHADE-Arena [35] for tool-calling agents and our new CUA-SHADE-Arena, which extends TheAgentCompany [60], for computer-use agents. We run MRT on existing LLM monitor scaffoldings, which orchestrate LLMs and parse agent trajectories, alongside a new hybrid hierarchical-sequential scaffolding proposed in this work. Our empirical results yield three key findings. First, agent awareness dominates monitor awareness: an agent's knowledge that it is being monitored substantially degrades the monitor's reliability. On the contrary, providing the monitor with more information about the agent is less helpful than expected. Second, monitor scaffolding matters more than monitor awareness: the hybrid scaffolding consistently outperforms baseline monitor scaffolding, and can enable weaker models to reliably monitor stronger agents -a weak-to-strong scaling effect. Third, in a human-in-the-loop setting where humans discuss with the LLM monitor to get an updated judgment for the agent's behavior, targeted human oversight is most effective; escalating only pre-flagged cases to human reviewers improved the TPR by approximately 15% at FPR = 0.01. Our work establishes a standard workflow for MRT, highlighting the lack of adversarial robustness for LLMs and humans when monitoring and detecting agent misbehavior. We release code, data, and logs to spur further research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6340e9d9-ee62-439d-a761-6be4f799c2b8Cited by top-tier papers5
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre et al.ICLR 2026 · 26 citations
- How does information access affect LLM monitors' ability to detect sabotage?Rauno Arike, Raja Moreno, Rohan Subramani, Shubhorup Biswas et al.ICML 2026 · 11 citations
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use AgentsJaylen Jones, Zhehao Zhang, Yuting Ning, Eric Fosler-Lussier et al.ICML 2026 · 9 citations
- AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk EvaluationChangyi Li, Pengfei Lu, Xudong Pan, Fazl Barez et al.ICML 2026 · 2 citations
- Peer-Preservation in Frontier ModelsYujin Potter, Nicholas Crispino, Vincent Siu, Chenguang Wang et al.ICML 2026
Builds on14
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 137 citations
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen et al.CCS 2024 · 132 citations
Related papers
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky et al.NeurIPS 2025 · 50 citations
- SafeSearch: Automated Red-Teaming of LLM-Based Search AgentsJianshuo Dong, Sheng Guo, Hao Wang, Xun Chen et al.ICML 2026 · 3 citations
- Constitutional Black-Box Monitoring for Scheming in LLM AgentsSimon Storf, Rich Barton-Cooper, James Peters-Gill, Marius HobbhahnICML 2026 · 1 citation
- Automated Red Teaming with GOAT: the Generative Offensive Agent TesterMaya Pavlova, Erik Brinkman, Krithika Iyer, Vítor Albiero et al.ICML 2025
- Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM AgentsYuchen Shao, Ziqun Bao, Yuheng Huang, Yuling Shi et al.ISSTA 2026
