The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)
Zihao Wang, Yibo Jiang, Jiahao Yu, Heqing Huang
Abstract
Large language models (LLMs) that integrate multiple input roles (e.g., system instructions, user queries, external tool outputs) are increasingly prevalent in practice. Ensuring that the model accurately distinguishes messages from each role -a concept we call role separationis crucial for consistent multi-role behavior. Although recent work often targets state-of-the-art prompt injection defenses, it remains unclear whether such methods truly teach LLMs to differentiate roles or merely memorize known triggers. In this paper, we examine role-separation learning: the process of teaching LLMs to robustly distinguish system and user tokens. Through a simple, controlled experimental framework, we find that finetuned models often rely on two proxies for role identification: (1) task type exploitation, and (2) proximity to begin-of-text. Although data augmentation can partially mitigate these shortcuts, it generally leads to iterative patching rather than a deeper fix. To address this, we propose reinforcing invariant signals that mark role boundaries by adjusting token-wise cues in the model's input encoding. In particular, manipulating position IDs helps the model learn clearer distinctions and reduces reliance on superficial proxies. By focusing on this mechanism-centered perspective, our work illuminates how LLMs can more reliably maintain consistent multi-role behavior without merely memorizing known prompts or triggers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a508cfa1-36d1-4444-8bfc-eefcbfea0c8eCited by top-tier papers3
- Prompt Injection as Role ConfusionCharles Ye, Jasmine Cui, Dylan Hadfield-MenellICML 2026 · 6 citations
- DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual FusionRuofan Liu, Yun Lin, Zhiyong Huang, Jin Song DongCCS 2026 · 3 citations
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMsTengyun Ma, Jiaqi Yao, Daojing He, Shihao Peng et al.NeurIPS 2025
Builds on6
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online GameSam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato et al.ICLR 2024 · 123 citations
- PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise TrainingDawei Zhu, Nan Yang, Liang Wang, Yifan Song et al.ICLR 2024 · 110 citations
- Instructional Segment Embedding: Improving LLM Safety with Instruction HierarchyTong Wu, Shujian Zhang, Kaiqiang Song, Silei Xu et al.ICLR 2025
Related papers
- Defenses Against Prompt Attacks Learn Surface HeuristicsShawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li et al.ACL 2026 · 8 citations
- ASIDE: Architectural Separation of Instructions and Data in Language ModelsEgor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova et al.ICLR 2026 · 28 citations
- Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?Egor Zverev, Sahar Abdelnabi, Soroush Tabesh, Mario Fritz et al.ICLR 2025
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 1 citation
- Control Illusion: The Failure of Instruction Hierarchies in Large Language ModelsYilin Geng, Haonan Li, Honglin Mu, Xudong Han et al.AAAI 2026 · 21 citations
