Adversaries Can Misuse Combinations of Safe Models
Erik Jones, Anca D. Dragan, Jacob Steinhardt
Abstract
Developers try to evaluate whether an AI system can be misused by adversaries before releasing it; for example, they might test whether a model enables cyberoffense, user manipulation, or bioterrorism. In this work, we show that individually testing models for misuse is inadequate; adversaries can misuse combinations of models even when each individual model is safe. The adversary accomplishes this by first decomposing tasks into subtasks, then solving each subtask with the best-suited model. For example, an adversary might solve challenging-but-benign subtasks with an aligned frontier model, and easy-but-malicious subtasks with a weaker misaligned model. We study two decomposition methods: manual decomposition where a human identifies a natural decomposition of a task, and automated decomposition where a weak model generates benign tasks for a frontier model to solve, then uses the solutions in-context to solve the original task. Using these decompositions, we empirically show that adversaries can create vulnerable code, explicit images, python scripts for hacking, and manipulative tweets at much higher rates with combinations of models than either individual model. Our work suggests that even perfectly-aligned frontier systems can enable misuse without ever producing malicious outputs, and that red-teaming efforts should extend beyond single models in isolation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2693f8b7-bcb1-4664-a212-71292cfc720aCited by top-tier papers5
- Estimating Worst-Case Frontier Risks of Open-Weight LLMsEric Wallace, Olivia Watkins, Miles Wang, Kai Chen et al.ICLR 2026 · 31 citations
- Large-scale online deanonymization with LLMsSimon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni et al.USENIX Security 2026 · 20 citations
- Monitoring Decomposition Attacks with Lightweight Sequential MonitorsYueh-Han Chen, Nitish Joshi, Yulin Chen, Maksym Andriushchenko et al.ICLR 2026 · 15 citations
- Distillation Robustifies UnlearningBruce W. Lee, Addie Foote, Alex Infanger, Leni Shor et al.NeurIPS 2025 · 15 citations
- Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded OutputsJackson Kaunismaa, John Hughes, Christina Q. Knight, Avery Griffin et al.ICLR 2026 · 7 citations
Builds on20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
Related papers
- CTRL-ALT-DECEIT Sabotage Evaluations for Automated AI R&DFrancis Ward, Teun van der Weij, Hanna Gábor, Sam Martin et al.NeurIPS 2025 · 11 citations
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak SusceptibilityBrendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng et al.EMNLP 2025 · 1 citation
- Breach By A Thousand Leaks: Unsafe Information Leakage in 'Safe' AI ResponsesDavid Glukhov, Ziwen Han, Ilia Shumailov, Vardan Papyan et al.ICLR 2025
- Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM AgentsNivya Talokar, Ayush Kumar Tarun, Murari Mandal, Maksym Andriushchenko et al.ICML 2026 · 2 citations
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 13 citations
