ACL2026
Knowing-but-Doing: Diagnosing and Defending Role-Play-Driven LLMs Jailbreaks via Moral Disengagement
Haiming Qin, Jianxun Lian, Qimin Zhong, Mingyang Zhou, Hao Liao, Naipeng Chao
摘要
Large Language Models (LLMs) are increasingly deployed in role-play scenarios, but their safety implications remain undercharacterized. We present an explanatory framework grounded in Bandura's Moral Disengagement theory and introduce a diagnostic benchmark (MD-Trace) for role-play jailbreaks. In our experiments, role-play improves safety behavior for benign personas while increasing unsafe compliance for malicious ones. We observe a Knowing-but-Doing failure in which models recognize safety risks in their thinking traces yet proceed to comply with harmful requests. Mechanism analysis suggests that Moral Justification is dominant, with Disregard of Consequences appearing as a secondary pattern. We compare multiple attack and defense methods and find that the diagnosis aligns with observed failure modes. Finally, we propose MD-Shield, an introspectionbased defense that reduces attack success while maintaining Role Fidelity. The source code is publicly available at https://github.com/ lavapapa/MoralJustify/ .