On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
Sarah Ball, Greg Gluch, Shafi Goldwasser, Frauke Kreuter, Omer Reingold, Guy N. Rothblum
摘要
With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with a focus on filters to prevent the generation of unsafe information. Two natural points of intervention are the filtering of the input prompt before it reaches the model, and filtering the output after generation. Our main results demonstrate computational challenges in filtering both prompts and outputs. First, we show that there exist LLMs for which there are no efficient input-prompt filters: adversarial prompts that elicit harmful behavior can be easily constructed, which are computationally indistinguishable from benign prompts for any efficient filter. Our second main result identifies a natural setting in which output filtering is computationally intractable. All of our separation results are under cryptographic hardness assumptions. In addition to these core findings, we also formalize and study relaxed mitigation approaches, demonstrating further computational barriers. We conclude that safety cannot be achieved by designing filters external to the LLM internals (architecture and weights); in particular, black-box access to the LLM will not suffice. Based on our technical results, we argue that an aligned AI system’s intelligence cannot be separated from its judgment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Bypassing Prompt Guards in Production with Controlled-Release PromptingJaiden Fairoze, Sanjam Garg, Keewoo Lee, Mingyuan WangUSENIX Security 2026 · 被引用 8 次
- Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model UnlearningPuning Yang, Junchi Yu, Qizhou Wang, Phil Torr 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Making Smart Contracts SmarterLoi Luu, Duc-Hiep Chu, Hrishi Olickel, Prateek Saxena 等CCS 2016 · 被引用 2,306 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityAndrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg 等ICML 2024 · 被引用 177 次
- Understanding Catastrophic Forgetting in Language Models via Implicit InferenceSuhas Kotha, Jacob Mitchell Springer, Aditi RaghunathanICLR 2024 · 被引用 131 次
相关 Paper
- Jailbreak Open-Sourced Large Language Models via Enforced DecodingHangfan Zhang, Zhimeng Guo, Huaisheng Zhu, Bochuan Cao 等ACL 2024
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski 等NeurIPS 2023 · 被引用 412 次
- Safety Alignment Can Be Not Superficial With Explicit Safety SignalsJianwei Li, Jung-Eun KimICML 2025
- Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak AttacksYingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng 等NDSS 2026 · 被引用 5 次
