HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine Translation
David Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loïc Barrault, Marta R. Costa-jussà
Abstract
Hallucinations in machine translation are translations that contain information completely unrelated to the input. Omissions are translations that do not include some of the input information. While both cases tend to be catastrophic errors undermining user trust, annotated data with these types of pathologies is extremely scarce and is limited to a few high-resource languages. In this work, we release an annotated dataset for the hallucination and omission phenomena covering 18 translation directions with varying resource levels and scripts. Our annotation covers different levels of partial and full hallucinations as well as omissions both at the sentence and at the word level. Additionally, we revisit previous methods for hallucination and omission detection, show that conclusions made based on a single language pair largely do not hold for a large-scale evaluation, and establish new solid baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language GenerationAdam BouyamournEMNLP 2023 · 9 citations
- Quantifying the Plausibility of Context Reliance in Neural Machine TranslationGabriele Sarti, Grzegorz Chrupala, Malvina Nissim, Arianna BisazzaICLR 2024 · 8 citations
- Word Alignment as Preference for Machine TranslationQiyu Wu, Masaaki Nagata, Zhongtao Miao, Yoshimasa TsuruokaEMNLP 2024 · 4 citations
- HAT: Hallucination Annotation for TranslationRajen Chatterjee, Xintong Li, Paisarn Charoenpornsawat, Allen LeeACL 2026
- M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine TranslationHao Wang, Linlong Xu, Heng Liu, Yangyang Liu et al.ACL 2026
Builds on7
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 138 citations
- Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the TransformerJavier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano et al.EMNLP 2022 · 18 citations
- Optimal Transport for Unsupervised Hallucination Detection in Neural Machine TranslationNuno Miguel Guerreiro, Pierre Colombo, Pablo Piantanida, André F. T. MartinsACL 2023 · 8 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
Related papers
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM HallucinationSaad Obaid ul Islam, Anne Lauscher, Goran GlavasEMNLP 2025
- HalLoc: Token-level Localization of Hallucinations for Vision Language ModelsEunkyu Park, Minyeong Kim, Gunhee KimCVPR 2025
- ANAH: Analytical Annotation of Hallucinations in Large Language ModelsZiwei Ji, Yuzhe Gu, Wenwei Zhang, Chengqi Lyu et al.ACL 2024 · 8 citations
- K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language ModelsJaehyung Seo, Heuiseok LimICLR 2025
- A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao et al.ACL 2022 · 194 citations
