Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
Taiming Lu, Philipp Koehn
摘要
This paper investigates the propagation of information in multilingual large language models (LLMs) and evaluates the efficacy of various unlearning methods. We demonstrate that misinformation, regardless of the language it is in, once introduced into these models through training data, can spread across different languages, compromising the integrity and reliability of the generated content. Our findings reveal that standard unlearning techniques, which typically focus on English data, are insufficient in mitigating the spread of fake content in multilingual contexts and could inadvertently reinforce misinformation across languages. We show that only by addressing misinformative responses in both English and the original language of the fake data we can effectively eliminate it for all languages. This underscores the critical need for comprehensive unlearning strategies that consider the multilingual nature of modern LLMs to enhance their safety and reliability across landscapes. Code and data is accessible here: https://github.com/TaiMingLu/learn-unlearn .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- BiMind: A Dual-Head Reasoning Model with Attention-Geometry Adapter for Incorrect Information DetectionZhongxing Zhang, Emily K. Vraga, Jisu Huh, Jaideep SrivastavaACL 2026 · 被引用 1 次
- Multilingual Unlearning in LLMs: Transfer, Dynamics, and ReversibilityChaoyi Xiang, Olga Ohrimenko, Benjamin Rubinstein, Lea FrermannICML 2026
它引用的顶会 Paper12
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning 等ICML 2022 · 被引用 520 次
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 被引用 378 次
相关 Paper
- To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMsZohaib Khan, Mustafa Dogan, Ifeoma Okoh, Pouya Sadeghi 等ACL 2026
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM HallucinationSaad Obaid ul Islam, Anne Lauscher, Goran GlavasEMNLP 2025
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 被引用 230 次
- Disinformation Capabilities of Large Language ModelsIvan Vykopal, Matús Pikuliak, Ivan Srba, Róbert Móro 等ACL 2024
- JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak AttacksMasahiro Kaneko, Ayana Niwa, Timothy BaldwinICLR 2026 · 被引用 6 次
