Control Tax: The Price of Keeping AI in Check
Mikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre, Samuel Albanie
Abstract
The rapid integration of agentic AI into high-stakes real-world applications requires robust oversight mechanisms. The emerging field of AI Control (AIC) aims to provide such an oversight mechanism, but practical adoption depends heavily on implementation overhead. To study this problem better, we introduce the notion of Control tax---the operational and financial cost of integrating control measures into AI pipelines. Our work makes three key contributions to the field of AIC: (1) we introduce a theoretical framework that quantifies the Control Tax and maps classifier performance to safety assurances; (2) we conduct comprehensive evaluations of state-of-the-art language models in adversarial settings, where attacker models insert subtle backdoors into code while monitoring models attempt to detect these vulnerabilities; and (3) we provide empirical financial cost estimates for control protocols and develop optimized monitoring strategies that balance safety and cost-effectiveness while accounting for practical constraints like auditing budgets. Our framework enables practitioners to make informed decisions by systematically connecting safety guarantees with their costs, advancing AIC through principled economic feasibility assessment across different deployment contexts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 892c4544-2e72-4245-a025-790a4bb5f3c2Cited by top-tier papers5
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky et al.NeurIPS 2025 · 50 citations
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre et al.ICLR 2026 · 26 citations
- Combining Cost Constrained Runtime Monitors for AI SafetyTim Tian Hua, James Baskerville, Henri Lemoine, Mia Hopman et al.NeurIPS 2025 · 12 citations
- How does information access affect LLM monitors' ability to detect sabotage?Rauno Arike, Raja Moreno, Rohan Subramani, Shubhorup Biswas et al.ICML 2026 · 11 citations
- Same Question, Different Lies: Cross-Context Consistency (C³) for Black-Box Sandbagging DetectionYulong Lin, Pablo Bernabeu-Pérez, Benjamin Arnav, Lennie Wells et al.ICML 2026
Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- Debating with More Persuasive LLMs Leads to More Truthful AnswersAkbir Khan, John Hughes, Dan Valentine, Laura Ruis et al.ICML 2024 · 244 citations
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code CompletionRoei Schuster, Congzheng Song, Eran Tromer, Vitaly ShmatikovUSENIX Security 2021 · 199 citations
Related papers
- Adaptive Deployment of Untrusted LLMs Reduces Distributed ThreatsJiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt et al.ICLR 2025
- Towards More Practical Threat Models in Artificial Intelligence SecurityKathrin Grosse, Lukas Bieringer, Tarek R. Besold, Alexandre AlahiUSENIX Security 2024 · 27 citations
- AgentBound: Securing Execution Boundaries of AI AgentsChristoph Bühler, Matteo Biagiola, Luca Di Grazia, Guido SalvaneschiFSE 2026 · 1 citation
- SafeScientist: Enhancing AI Scientist Safety for Risk-Aware Scientific DiscoveryKunlun Zhu, Jiaxun Zhang, Ziheng Qi, Nuoxing Shang et al.EMNLP 2025
- Cloak, Honey, Trap: Proactive Defenses Against LLM AgentsDaniel Ayzenshteyn, Roy Weiss, Yisroel MirskyUSENIX Security 2025
