Control Tax: The Price of Keeping AI in Check
Mikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre, Samuel Albanie
摘要
The rapid integration of agentic AI into high-stakes real-world applications requires robust oversight mechanisms. The emerging field of AI Control (AIC) aims to provide such an oversight mechanism, but practical adoption depends heavily on implementation overhead. To study this problem better, we introduce the notion of Control tax---the operational and financial cost of integrating control measures into AI pipelines. Our work makes three key contributions to the field of AIC: (1) we introduce a theoretical framework that quantifies the Control Tax and maps classifier performance to safety assurances; (2) we conduct comprehensive evaluations of state-of-the-art language models in adversarial settings, where attacker models insert subtle backdoors into code while monitoring models attempt to detect these vulnerabilities; and (3) we provide empirical financial cost estimates for control protocols and develop optimized monitoring strategies that balance safety and cost-effectiveness while accounting for practical constraints like auditing budgets. Our framework enables practitioners to make informed decisions by systematically connecting safety guarantees with their costs, advancing AIC through principled economic feasibility assessment across different deployment contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky 等NeurIPS 2025 · 被引用 50 次
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre 等ICLR 2026 · 被引用 26 次
- Combining Cost Constrained Runtime Monitors for AI SafetyTim Tian Hua, James Baskerville, Henri Lemoine, Mia Hopman 等NeurIPS 2025 · 被引用 12 次
- How does information access affect LLM monitors' ability to detect sabotage?Rauno Arike, Raja Moreno, Rohan Subramani, Shubhorup Biswas 等ICML 2026 · 被引用 11 次
- Same Question, Different Lies: Cross-Context Consistency (C³) for Black-Box Sandbagging DetectionYulong Lin, Pablo Bernabeu-Pérez, Benjamin Arnav, Lennie Wells 等ICML 2026
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud 等ICLR 2024 · 被引用 762 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- Debating with More Persuasive LLMs Leads to More Truthful AnswersAkbir Khan, John Hughes, Dan Valentine, Laura Ruis 等ICML 2024 · 被引用 244 次
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code CompletionRoei Schuster, Congzheng Song, Eran Tromer, Vitaly ShmatikovUSENIX Security 2021 · 被引用 199 次
相关 Paper
- Adaptive Deployment of Untrusted LLMs Reduces Distributed ThreatsJiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt 等ICLR 2025
- Towards More Practical Threat Models in Artificial Intelligence SecurityKathrin Grosse, Lukas Bieringer, Tarek R. Besold, Alexandre AlahiUSENIX Security 2024 · 被引用 27 次
- AgentBound: Securing Execution Boundaries of AI AgentsChristoph Bühler, Matteo Biagiola, Luca Di Grazia, Guido SalvaneschiFSE 2026 · 被引用 1 次
- SafeScientist: Enhancing AI Scientist Safety for Risk-Aware Scientific DiscoveryKunlun Zhu, Jiaxun Zhang, Ziheng Qi, Nuoxing Shang 等EMNLP 2025
- Cloak, Honey, Trap: Proactive Defenses Against LLM AgentsDaniel Ayzenshteyn, Roy Weiss, Yisroel MirskyUSENIX Security 2025
