AI Control: Improving Safety Despite Intentional Subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien Roger
摘要
As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or redteaming techniques to surface subtle failure modes. However, researchers have not evaluated whether such techniques still ensure safety if the model is itself intentionally trying to subvert them. In this paper, we develop and evaluate pipelines of safety techniques ("protocols") that are robust to intentional subversion. We investigate a scenario in which we want to solve a sequence of programming problems, using access to a powerful but untrusted model (in our case, GPT-4), access to a less powerful trusted model (in our case, GPT-3.5), and limited access to high-quality trusted labor. We investigate protocols that aim to never submit solutions containing backdoors, which we operationalize here as logical errors that are not caught by test cases. We investigate a range of protocols and test each against strategies that the untrusted model could use to subvert them. One protocol is what we call trusted editing. This protocol first asks GPT-4 to write code, and then asks GPT-3.5 to rate the suspiciousness of that code. If the code is below some suspiciousness threshold, it is submitted. Otherwise, GPT-3.5 edits the solution to remove parts that seem suspicious and then submits the edited code. Another protocol is untrusted monitoring. This protocol asks GPT-4 to write code, and then asks another instance of GPT-4 whether the code is backdoored, using various techniques to prevent the GPT-4 instances from colluding. These protocols improve substantially on simple baselines. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Debating with More Persuasive LLMs Leads to More Truthful AnswersAkbir Khan, John Hughes, Dan Valentine, Laura Ruis 等ICML 2024 · 被引用 244 次
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina 等NeurIPS 2024 · 被引用 140 次
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes 等NeurIPS 2025 · 被引用 52 次
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky 等NeurIPS 2025 · 被引用 50 次
- Large language models can learn and generalize steganographic chain-of-thought under process supervisionRobert MC Carthy, Joey Skaf, Luis Ibañez-Lissen, Vasil Georgiev 等NeurIPS 2025 · 被引用 29 次
它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu 等ICLR 2024 · 被引用 202 次
相关 Paper
- Adaptive Deployment of Untrusted LLMs Reduces Distributed ThreatsJiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt 等ICLR 2025
- Instruction Backdoor Attacks Against Customized LLMsRui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang 等USENIX Security 2024 · 被引用 83 次
- CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional SafetyLu Yan, Zhuo Zhang, Xiangzhe Xu, Shengwei An 等ISSTA 2026
- Simulate and Eliminate: Revoke Backdoors for Generative Large Language ModelsHaoran Li, Yulin Chen, Zihao Zheng, Qi Hu 等AAAI 2025 · 被引用 6 次
- AI Sandbagging: Language Models can Strategically Underperform on EvaluationsTeun van der Weij, Felix Hofstätter, Oliver Jaffe, Samuel F. Brown 等ICLR 2025
