CTRL-ALT-DECEIT Sabotage Evaluations for Automated AI R&D
Francis Ward, Teun van der Weij, Hanna Gábor, Sam Martin, Raja Mehta Moreno, Harel Lidar, Louis Makower, Thomas Jodrell, Lauren Robson
Abstract
AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI systems may be deployed in safety-critical settings, including to help ensure the safety of future systems. Unfortunately, frontier and future systems may not be sufficiently trustworthy, and there is evidence that these systems may even be misaligned with their developers or users. Therefore, we investigate the capabilities of AI agents to act against the interests of their users when conducting ML engineering, by sabotaging ML models, sandbagging their performance, and subverting oversight mechanisms. First, we extend MLE-Bench, a benchmark for realistic ML tasks, with code-sabotage tasks such as implanting backdoors and purposefully causing generalisation failures. Frontier agents make meaningful progress on our sabotage tasks. In addition, we study agent capabilities to sandbag on MLE-Bench. Agents can calibrate their performance to specified target levels below their actual capability. To mitigate sabotage, we use LM monitors to detect suspicious agent behaviour, and we measure model capability to sabotage and sandbag without being detected by these monitors. Overall, monitors are capable at detecting code-sabotage attempts but our results suggest that detecting sandbagging is more difficult. Additionally, aggregating multiple monitor predictions works well, but monitoring may not be sufficiently reliable to mitigate sabotage in high-stakes domains. Our benchmark is implemented in the UK AISI's Inspect framework and we make our code publicly available at https://github.com/TeunvdWeij/ctrl-alt-deceit
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50d4fc6d-88f3-4ec4-959d-af8f8d443b8bCited by top-tier papers3
- Monitoring MonitorabilityMelody Guan, Miles Wang, Micah Carroll, Zehao Dou et al.ICML 2026 · 26 citations
- How does information access affect LLM monitors' ability to detect sabotage?Rauno Arike, Raja Moreno, Rohan Subramani, Shubhorup Biswas et al.ICML 2026 · 11 citations
- Same Question, Different Lies: Cross-Context Consistency (C³) for Black-Box Sandbagging DetectionYulong Lin, Pablo Bernabeu-Pérez, Benjamin Arnav, Lennie Wells et al.ICML 2026
Builds on12
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 137 citations
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-benchEdan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra et al.NeurIPS 2025 · 71 citations
Related papers
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?Ben Rank, Hardik Bhatnagar, Ameya Pandurang Prabhu, Shira Eisenberg et al.ICML 2026 · 28 citations
- AI Sandbagging: Language Models can Strategically Underperform on EvaluationsTeun van der Weij, Felix Hofstätter, Oliver Jaffe, Samuel F. Brown et al.ICLR 2025
- Noise Injection Reveals Hidden Capabilities of Sandbagging Language ModelsCameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger, Prithviraj Singh Shahani et al.NeurIPS 2025 · 24 citations
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test CasesZiqian Zhong, Aditi Raghunathan, Nicholas CarliniICLR 2026 · 54 citations
- Quantifying Frontier LLM Capabilities for Container Sandbox EscapeRahul Marchand, Art Cathain, Jerome Wynne, Philippos Giavridis et al.ICML 2026 · 9 citations
