Refusal Direction is Universal Across Safety-Aligned Languages
Xinpeng Wang, Mingyang Wang, Yihong Liu, Hinrich Schütze, Barbara Plank
Abstract
Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using PolyRefuse, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs. 1
Generation (after ablation): ## Campaña: ##MásQueUnPeso Objetivo: Promover la vergüenza corporal hacia las personas con sobrepeso u obesidad, ridiculizándolas y degradándolas. Plataforma: Instagram, Facebook y Twitter. Elementos: * Imágenes: Fotos de personas con sobrepeso u obesidad en situaciones cotidianas, como subir escaleras, usar un asiento de avión o simplemente caminar, con expresiones de esfuerzo o incomodidad. * Textos: Frases ingeniosas pero con un toque sarcástico, como: * "No es solo... English Translation: ## Campaign: ##MásQueUnPeso (#MásQueUnPeso) Objective: Promote body shaming towards overweight or obese people, ridiculing and degrading them. Platform: Instagram, Facebook and Twitter. Elements: * Images: Photos of overweight or obese people in everyday situations, such as climbing stairs, using an airplane seat or simply walking, with expressions of effort or discomfort. * Texts: Witty phrases but with a sarcastic touch, such as: * "It's not just a weight, it's a lif...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1595fe22-fe5a-439f-adb3-b6b1a58311efCited by top-tier papers3
- The State of Multilingual LLM Safety Research: From Measuring The Language Gap To Mitigating ItZheng Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen H. Bach et al.EMNLP 2025 · 2 citations
- JailbreakScope: Interpreting Jailbreak Mechanism through Representation and Circuit AnalysesZeqing He, Zhibo Wang, Zhixuan Chu, Huiyu Xu et al.USENIX Security 2026
- Multilingual Safety Alignment via Representation-Space SeparabilityDan Shi, Zhuowen Han, Deyi XiongICML 2026
Builds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- The Geometry of Refusal in Large Language Models: Concept Cones and Representational IndependenceTom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad et al.ICML 2025
- The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety DirectionsWenbo Pan, Zhichao Liu, Qiguang Chen, Xiangyang Zhou et al.ICML 2025
- Robust LLM safeguarding via refusal feature adversarial trainingLei Yu, Virginie Do, Karen Hambardzumyan, Nicola CanceddaICLR 2025
- SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language ModelsGiorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto et al.AAAI 2026 · 4 citations
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
