Detecting and Mitigating Hallucinations in Multilingual Summarisation
Yifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen
摘要
Hallucinations pose a significant challenge to the reliability of neural models for abstractive summarisation. While automatically generated summaries may be fluent, they often lack faithfulness to the original document. This issue becomes even more pronounced in low-resource languages, where summarisation requires cross-lingual transfer. With the existing faithful metrics focusing on English, even measuring the extent of this phenomenon in cross-lingual settings is hard. To address this, we first develop a novel metric, mFACT, evaluating the faithfulness of non-English summaries, leveraging translation-based transfer from multiple English faithfulness metrics. Through extensive experiments in multiple languages, we demonstrate that mFACT is best suited to detect hallucinations compared to alternative metrics. With mFACT, we assess a broad range of multilingual large language models, and find that they all tend to hallucinate often in languages different from English. We then propose a simple but effective method to reduce hallucinations in cross-lingual transfer, which weighs the loss of each training example by its faithfulness score. This method drastically increases both performance and faithfulness according to both automatic and human evaluation when compared to strong baselines for cross-lingual transfer such as MAD-X. Our code and dataset are available at https://github.com/yfqiu-nlp/mfact-summ.<br/>
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Spectral Editing of Activations for Large Language Model AlignmentYifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen 等NeurIPS 2024 · 被引用 66 次
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViTGuy Bar-Shalom, Fabrizio Frasca, Yaniv Galron, Yftah Ziser 等NeurIPS 2025 · 被引用 17 次
- An Audit on the Perspectives and Challenges of Hallucinations in NLPPranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs 等EMNLP 2024 · 被引用 8 次
- Beyond Next Token Probabilities: Learnable, Fast Detection of Hallucinations and Data Contamination on LLM Output DistributionsGuy Bar-Shalom, Fabrizio Frasca, Derek Lim, Yoav Gelberg 等AAAI 2026 · 被引用 7 次
- Neural Message-Passing on Attention Graphs for Hallucination DetectionFabrizio Frasca, Guy Bar-Shalom, Yftah Ziser, Haggai MaronICLR 2026 · 被引用 6 次
它引用的顶会 Paper20
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 被引用 329 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- Knowledge Graph-Augmented Abstractive Summarization with Semantic-Driven Cloze RewardLuyang Huang, Lingfei Wu, Lu WangACL 2020 · 被引用 152 次
- SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive SummarizationMathieu Ravaut, Shafiq R. Joty, Nancy F. ChenACL 2022 · 被引用 116 次
相关 Paper
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM HallucinationSaad Obaid ul Islam, Anne Lauscher, Goran GlavasEMNLP 2025
- CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality EvaluationYexing Du, Kaiyuan Liu, Youcheng Pan, Zheng Chu 等AAAI 2026 · 被引用 4 次
- Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language ModelQi Jia, Siyu Ren, Yizhu Liu, Kenny Q. ZhuEMNLP 2023 · 被引用 4 次
- MultiSumm: Towards a Unified Model for Multi-Lingual Abstractive SummarizationYue Cao, Xiaojun Wan, Jin-ge Yao, Dian YuAAAI 2020 · 被引用 28 次
- HAT: Hallucination Annotation for TranslationRajen Chatterjee, Xintong Li, Paisarn Charoenpornsawat, Allen LeeACL 2026
