HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine Translation
David Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loïc Barrault, Marta R. Costa-jussà
摘要
Hallucinations in machine translation are translations that contain information completely unrelated to the input. Omissions are translations that do not include some of the input information. While both cases tend to be catastrophic errors undermining user trust, annotated data with these types of pathologies is extremely scarce and is limited to a few high-resource languages. In this work, we release an annotated dataset for the hallucination and omission phenomena covering 18 translation directions with varying resource levels and scripts. Our annotation covers different levels of partial and full hallucinations as well as omissions both at the sentence and at the word level. Additionally, we revisit previous methods for hallucination and omission detection, show that conclusions made based on a single language pair largely do not hold for a large-scale evaluation, and establish new solid baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language GenerationAdam BouyamournEMNLP 2023 · 被引用 9 次
- Quantifying the Plausibility of Context Reliance in Neural Machine TranslationGabriele Sarti, Grzegorz Chrupala, Malvina Nissim, Arianna BisazzaICLR 2024 · 被引用 8 次
- Word Alignment as Preference for Machine TranslationQiyu Wu, Masaaki Nagata, Zhongtao Miao, Yoshimasa TsuruokaEMNLP 2024 · 被引用 4 次
- HAT: Hallucination Annotation for TranslationRajen Chatterjee, Xintong Li, Paisarn Charoenpornsawat, Allen LeeACL 2026
- M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine TranslationHao Wang, Linlong Xu, Heng Liu, Yangyang Liu 等ACL 2026
它引用的顶会 Paper7
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 被引用 138 次
- Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the TransformerJavier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano 等EMNLP 2022 · 被引用 18 次
- Optimal Transport for Unsupervised Hallucination Detection in Neural Machine TranslationNuno Miguel Guerreiro, Pierre Colombo, Pablo Piantanida, André F. T. MartinsACL 2023 · 被引用 8 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan 等ACL 2022
相关 Paper
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM HallucinationSaad Obaid ul Islam, Anne Lauscher, Goran GlavasEMNLP 2025
- HalLoc: Token-level Localization of Hallucinations for Vision Language ModelsEunkyu Park, Minyeong Kim, Gunhee KimCVPR 2025
- ANAH: Analytical Annotation of Hallucinations in Large Language ModelsZiwei Ji, Yuzhe Gu, Wenwei Zhang, Chengqi Lyu 等ACL 2024 · 被引用 8 次
- K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language ModelsJaehyung Seo, Heuiseok LimICLR 2025
- A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao 等ACL 2022 · 被引用 194 次
