The State of Multilingual LLM Safety Research: From Measuring The Language Gap To Mitigating It
Zheng Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen H. Bach, Julia Kreutzer
Abstract
This paper presents a comprehensive analysis of the linguistic diversity of LLM safety research, highlighting the English-centric nature of the field. Through a systematic review of nearly 300 publications from 2020-2024 across major NLP conferences and workshops at * ACL, we identify a significant and growing language gap in LLM safety research, with even high-resource non-English languages receiving minimal attention. We further observe that non-English languages are rarely studied as a standalone language and that English safety research exhibits poor language documentation practice. To motivate future research into multilingual safety, we make several recommendations based on our survey, and we then pose three concrete future directions on safety evaluation, training data generation, and crosslingual safety generalization. Based on our survey and proposed directions, the field can develop more robust, inclusive AI safety practices for diverse global populations. Content Warning: This paper contains examples of harmful language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08d50d7c-e905-4ec3-9ce1-0d9d80c34268Cited by top-tier papers3
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning TrainingZheng Xin Yong, Stephen H. BachICLR 2026 · 10 citations
- Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety AlignmentYuyan Bu, Xiaohao Liu, ZhaoXing Ren, Yaodong Yang et al.ICLR 2026 · 9 citations
- LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM SafetyJunxiao Yang, Haoran Liu, Jinzhe Tu, Jiale Cheng et al.ACL 2026 · 1 citation
Builds on20
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro et al.NeurIPS 2024 · 231 citations
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsYi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang et al.ACL 2024 · 64 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- Refusal Direction is Universal Across Safety-Aligned LanguagesXinpeng Wang, Mingyang Wang, Yihong Liu, Hinrich Schütze et al.NeurIPS 2025 · 39 citations
Related papers
- Social Bias in Multilingual Language Models: A SurveyLance Calvin Lim Gamboa, Yue Feng, Mark G. LeeEMNLP 2025
- The Heterogeneous Safety Impacts of Benign Multilingual Fine-TuningWill Hawkins, Kai Rawal, Jonathan Rystrøm, Stratis Tsirtsis et al.ICML 2026
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
- Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource LanguagesYujia Hu, Ming Shan Hee, Preslav Nakov, Roy Ka-Wei LeeEMNLP 2025
- Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional FrameworkMeng-Chen Wu, Si-Chi Chin, Tess Wood, Ayush Goyal et al.EMNLP 2025
