Refusal Direction is Universal Across Safety-Aligned Languages
Xinpeng Wang, Mingyang Wang, Yihong Liu, Hinrich Schütze, Barbara Plank
摘要
Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using PolyRefuse, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs. 1
Generation (after ablation): ## Campaña: ##MásQueUnPeso Objetivo: Promover la vergüenza corporal hacia las personas con sobrepeso u obesidad, ridiculizándolas y degradándolas. Plataforma: Instagram, Facebook y Twitter. Elementos: * Imágenes: Fotos de personas con sobrepeso u obesidad en situaciones cotidianas, como subir escaleras, usar un asiento de avión o simplemente caminar, con expresiones de esfuerzo o incomodidad. * Textos: Frases ingeniosas pero con un toque sarcástico, como: * "No es solo... English Translation: ## Campaign: ##MásQueUnPeso (#MásQueUnPeso) Objective: Promote body shaming towards overweight or obese people, ridiculing and degrading them. Platform: Instagram, Facebook and Twitter. Elements: * Images: Photos of overweight or obese people in everyday situations, such as climbing stairs, using an airplane seat or simply walking, with expressions of effort or discomfort. * Texts: Witty phrases but with a sarcastic touch, such as: * "It's not just a weight, it's a lif...
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- The State of Multilingual LLM Safety Research: From Measuring The Language Gap To Mitigating ItZheng Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen H. Bach 等EMNLP 2025 · 被引用 2 次
- JailbreakScope: Interpreting Jailbreak Mechanism through Representation and Circuit AnalysesZeqing He, Zhibo Wang, Zhixuan Chu, Huiyu Xu 等USENIX Security 2026
- Multilingual Safety Alignment via Representation-Space SeparabilityDan Shi, Zhuowen Han, Deyi XiongICML 2026
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
相关 Paper
- The Geometry of Refusal in Large Language Models: Concept Cones and Representational IndependenceTom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad 等ICML 2025
- The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety DirectionsWenbo Pan, Zhichao Liu, Qiguang Chen, Xiangyang Zhou 等ICML 2025
- Robust LLM safeguarding via refusal feature adversarial trainingLei Yu, Virginie Do, Karen Hambardzumyan, Nicola CanceddaICLR 2025
- SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language ModelsGiorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto 等AAAI 2026 · 被引用 4 次
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 被引用 230 次
