ASIDE: Architectural Separation of Instructions and Data in Language Models
Egor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova, Rush Tabesh, Sebastian Lapuschkin, Wojciech Samek, Christoph H. Lampert
摘要
Despite their remarkable performance, large language models lack elementary safety features, making them susceptible to numerous malicious attacks. In particular, previous work has identified the absence of an intrinsic separation between instructions and data as the root cause of the success of prompt injection attacks. In this work, we propose a new architectural element, ASIDE, that allows language models to clearly separate instructions and data at the level of token embeddings. ASIDE applies an orthogonal rotation to the embeddings of data tokens, thus creating clearly distinct representations of instructions and data tokens without introducing any additional parameters. As we demonstrate experimentally across a range of models, instruction-tuning LLMs with ASIDE (1) achieves substantially higher instruction-data separation without performance loss and (2) makes the models more robust to prompt injection benchmarks, even without dedicated safety training. Additionally, we provide insights into the mechanism underlying our method through an analysis of the model representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre 等ICLR 2026 · 被引用 26 次
- Prompt Injection as Role ConfusionCharles Ye, Jasmine Cui, Dylan Hadfield-MenellICML 2026 · 被引用 6 次
- Beyond Oracle: Verifier-Supervision for Instruction Hierarchy in Reasoning and Instruction-Tuned LLMsSian-Yao Huang, Li-Hsien Chang, Che-Yu Lin, Cheng-Lin YangNeurIPS 2025 · 被引用 4 次
- DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual FusionRuofan Liu, Yun Lin, Zhiyong Huang, Jin Song DongCCS 2026 · 被引用 3 次
- Don't Forget the Enjoin: FocalLoRA for Instruction Hierarchical Alignment in Large Language ModelsZitong Shi, Frank Wan, Haixin Wang, Ruoyan Li 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper18
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas 等NeurIPS 2024 · 被引用 362 次
相关 Paper
- Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?Egor Zverev, Sahar Abdelnabi, Soroush Tabesh, Mario Fritz 等ICLR 2025
- StruQ: Defending Against Prompt Injection with Structured QueriesSizhe Chen, Julien Piet, Chawin Sitawarin, David A. WagnerUSENIX Security 2025
- Instructional Segment Embedding: Improving LLM Safety with Instruction HierarchyTong Wu, Shujian Zhang, Kaiqiang Song, Silei Xu 等ICLR 2025
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt InjectionZekun Li, Baolin Peng, Pengcheng He, Xifeng YanEMNLP 2024 · 被引用 15 次
- The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)Zihao Wang, Yibo Jiang, Jiahao Yu, Heqing HuangICML 2025
