Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
Taiming Lu, Philipp Koehn
Abstract
This paper investigates the propagation of information in multilingual large language models (LLMs) and evaluates the efficacy of various unlearning methods. We demonstrate that misinformation, regardless of the language it is in, once introduced into these models through training data, can spread across different languages, compromising the integrity and reliability of the generated content. Our findings reveal that standard unlearning techniques, which typically focus on English data, are insufficient in mitigating the spread of fake content in multilingual contexts and could inadvertently reinforce misinformation across languages. We show that only by addressing misinformative responses in both English and the original language of the fake data we can effectively eliminate it for all languages. This underscores the critical need for comprehensive unlearning strategies that consider the multilingual nature of modern LLMs to enhance their safety and reliability across landscapes. Code and data is accessible here: https://github.com/TaiMingLu/learn-unlearn .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e4aeb4b-dc5c-4465-a207-41b347d7a1b5Cited by top-tier papers2
- BiMind: A Dual-Head Reasoning Model with Attention-Geometry Adapter for Incorrect Information DetectionZhongxing Zhang, Emily K. Vraga, Jisu Huh, Jaideep SrivastavaACL 2026 · 1 citation
- Multilingual Unlearning in LLMs: Transfer, Dynamics, and ReversibilityChaoyi Xiang, Olga Ohrimenko, Benjamin Rubinstein, Lea FrermannICML 2026
Builds on12
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning et al.ICML 2022 · 520 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
Related papers
- To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMsZohaib Khan, Mustafa Dogan, Ifeoma Okoh, Pouya Sadeghi et al.ACL 2026
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM HallucinationSaad Obaid ul Islam, Anne Lauscher, Goran GlavasEMNLP 2025
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
- Disinformation Capabilities of Large Language ModelsIvan Vykopal, Matús Pikuliak, Ivan Srba, Róbert Móro et al.ACL 2024
- JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak AttacksMasahiro Kaneko, Ayana Niwa, Timothy BaldwinICLR 2026 · 6 citations
