To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs
Zohaib Khan, Mustafa Dogan, Ifeoma Okoh, Pouya Sadeghi, Siddhartha Shrestha, Sergius Justus Chesami Nyah, Mahmoud O. Mokhiamar, Michael J. Ryan, Tarek Naous
摘要
Misinformation is on the rise, and the strong writing capabilities of LLMs lower the barrier for malicious actors to produce and disseminate false information. We study how LLMs behave when prompted to spread misinformation across languages and target countries, and introduce GlobalLies, a multilingual parallel dataset of 440 misinformation generation prompt templates and 6,867 entities, spanning 8 languages and 195 countries. Using both human annotations and large-scale LLM-as-a-judge evaluations across hundreds of thousands of generations from state-of-the-art models, we show that misinformation generation varies systematically based on the country being discussed. Propagation of lies by LLMs is substantially higher in many lower-resource languages and for countries with a lower Human Development Index (HDI). We find that existing mitigation strategies provide uneven protection: input safety classifiers exhibit cross-lingual gaps, and retrieval-augmented fact-checking remains inconsistent across regions due to unequal information availability. We release GlobalLies for research purposes, aiming to support the development of mitigation strategies to reduce the spread of global misinformation: https://github.com/zohaib-khan5040/globallies
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 被引用 230 次
- Stanceosaurus: Classifying Stance Towards Multicultural MisinformationJonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu 等EMNLP 2022 · 被引用 6 次
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun 等ACL 2024
- Disinformation Capabilities of Large Language ModelsIvan Vykopal, Matús Pikuliak, Ivan Srba, Róbert Móro 等ACL 2024
- Analyzing Leakage of Personally Identifiable Information in Language ModelsNils Lukas, Ahmed Salem, Robert Sim, Shruti Tople 等S&P 2023
相关 Paper
- Learn and Unlearn: Addressing Misinformation in Multilingual LLMsTaiming Lu, Philipp KoehnEMNLP 2025
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM HallucinationSaad Obaid ul Islam, Anne Lauscher, Goran GlavasEMNLP 2025
- SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful RequestsPunya Syon Pandey, Hai Son Le, Devansh Bhardwaj, Zhijing JinICLR 2026 · 被引用 7 次
- JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak AttacksMasahiro Kaneko, Ayana Niwa, Timothy BaldwinICLR 2026 · 被引用 6 次
- Truth Knows No Language: Evaluating Truthfulness Beyond EnglishBlanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes 等ACL 2025
