Adversaries Can Misuse Combinations of Safe Models
Erik Jones, Anca D. Dragan, Jacob Steinhardt
摘要
Developers try to evaluate whether an AI system can be misused by adversaries before releasing it; for example, they might test whether a model enables cyberoffense, user manipulation, or bioterrorism. In this work, we show that individually testing models for misuse is inadequate; adversaries can misuse combinations of models even when each individual model is safe. The adversary accomplishes this by first decomposing tasks into subtasks, then solving each subtask with the best-suited model. For example, an adversary might solve challenging-but-benign subtasks with an aligned frontier model, and easy-but-malicious subtasks with a weaker misaligned model. We study two decomposition methods: manual decomposition where a human identifies a natural decomposition of a task, and automated decomposition where a weak model generates benign tasks for a frontier model to solve, then uses the solutions in-context to solve the original task. Using these decompositions, we empirically show that adversaries can create vulnerable code, explicit images, python scripts for hacking, and manipulative tweets at much higher rates with combinations of models than either individual model. Our work suggests that even perfectly-aligned frontier systems can enable misuse without ever producing malicious outputs, and that red-teaming efforts should extend beyond single models in isolation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Estimating Worst-Case Frontier Risks of Open-Weight LLMsEric Wallace, Olivia Watkins, Miles Wang, Kai Chen 等ICLR 2026 · 被引用 31 次
- Large-scale online deanonymization with LLMsSimon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni 等USENIX Security 2026 · 被引用 20 次
- Monitoring Decomposition Attacks with Lightweight Sequential MonitorsYueh-Han Chen, Nitish Joshi, Yulin Chen, Maksym Andriushchenko 等ICLR 2026 · 被引用 15 次
- Distillation Robustifies UnlearningBruce W. Lee, Addie Foote, Alex Infanger, Leni Shor 等NeurIPS 2025 · 被引用 15 次
- Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded OutputsJackson Kaunismaa, John Hughes, Christina Q. Knight, Avery Griffin 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
相关 Paper
- CTRL-ALT-DECEIT Sabotage Evaluations for Automated AI R&DFrancis Ward, Teun van der Weij, Hanna Gábor, Sam Martin 等NeurIPS 2025 · 被引用 11 次
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak SusceptibilityBrendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng 等EMNLP 2025 · 被引用 1 次
- Breach By A Thousand Leaks: Unsafe Information Leakage in 'Safe' AI ResponsesDavid Glukhov, Ziwen Han, Ilia Shumailov, Vardan Papyan 等ICLR 2025
- Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM AgentsNivya Talokar, Ayush Kumar Tarun, Murari Mandal, Maksym Andriushchenko 等ICML 2026 · 被引用 2 次
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 被引用 13 次
