Multilingual Safety Alignment Via Sparse Weight Editing
Jiaming Liang, Zhaoxin Wang, Handing Wang
Abstract
Large Language Models (LLMs) exhibit significant safety disparities across languages, with low-resource languages (LRLs) often bypassing safety guardrails established for high-resource languages (HRLs) like English. Existing solutions, such as multilingual supervised fine-tuning (SFT) or Reinforcement Learning from Human Feedback (RLHF), are computationally expensive and de- pendent on scarce multilingual safety data. In this paper, we propose a novel, training-free alignment framework based on Sparse Weight Editing. Identifying that safety capabilities are localized within a sparse set of ”safety neurons,” we formulate the cross-lingual alignment problem as a constrained linear transformation. We derive a closed-form solution to optimally map the harmful representations of LRLs to the robust safety subspaces of HRLs, while preserving general utility via a null-space projection constraint. Extensive experiments across 8 languages and multiple model families (Llama-3, Qwen-2.5) demonstrate that our method significantly reduces Attack Success Rate (ASR) in LRLs with negligible impact on general reasoning capabilities, all achieved with a single, data-efficient calculation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7a7feca-e5a8-4b9a-9273-4fa3723fa286Builds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- Many-shot JailbreakingCem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma et al.NeurIPS 2024 · 338 citations
Related papers
- Multilingual Safety Alignment via Representation-Space SeparabilityDan Shi, Zhuowen Han, Deyi XiongICML 2026
- Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety NeuronsXianhui Zhang, Chengyu Xie, Linxia Zhu, Yonghui Yang et al.ICML 2026 · 5 citations
- LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM SafetyJunxiao Yang, Haoran Liu, Jinzhe Tu, Jiale Cheng et al.ACL 2026 · 1 citation
- NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-TuningXin Yi, Shunfan Zheng, Linlin Wang, Gerard de Melo et al.AAAI 2025 · 38 citations
- MPO: Multilingual Safety Alignment via Reward Gap OptimizationWeixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu et al.ACL 2025 · 12 citations
