SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question Answering
HyoJung Han, Marine Carpuat, Jordan L. Boyd-Graber
Abstract
Detractors of neural machine translation admit that while its translations are fluent, it sometimes gets key facts wrong. This is particularly important in simultaneous interpretation where translations have to be provided as fast as possible: before a sentence is complete. Yet, evaluations of simultaneous machine translation (SimulMT) fail to capture if systems correctly translate the most salient elements of a question: people, places, and dates. To address this problem, we introduce a downstream word-by-word question answering evaluation task (SimQA): given a source language question, translate the question word by word into the target language, and answer as soon as possible. SimQA jointly measures whether the SimulMT models translate the question quickly and accurately, and can reveal shortcomings in existing neural systems—hallucinating or omitting facts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0da4d7b4-5f6d-4f2f-87fd-26d57f857098Cited by top-tier papers4
- Bridging Background Knowledge Gaps in Translation with Automatic ExplicitationHyoJung Han, Jordan L. Boyd-Graber, Marine CarpuatEMNLP 2023 · 3 citations
- LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question AnsweringRan Zhang, Wei Zhao, Lieve Macken, Steffen EgerEMNLP 2025 · 2 citations
- Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine TranslationDayeon Ki, Kevin Duh, Marine CarpuatEMNLP 2025
- Measuring User's Mental Models of Speech Translation in Human-AI CollaborationHyojung Han, Nishant Balepur, Jordan Lee Boyd-Graber, Marine CarpuatACL 2026
Builds on9
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 90 citations
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu et al.EMNLP 2020 · 41 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- Simul-LLM: A Framework for Exploring High-Quality Simultaneous Translation with Large Language ModelsVictor Agostinelli, Max Wild, Matthew Raffel, Kazi Ahmed Asif Fuad et al.ACL 2024 · 4 citations
- Extrinsic Evaluation of Machine Translation MetricsNikita Moghe, Tom Sherborne, Mark Steedman, Alexandra BirchACL 2023 · 12 citations
- Improving Simultaneous Machine Translation with Monolingual DataHexuan Deng, Liang Ding, Xuebo Liu, Meishan Zhang et al.AAAI 2023 · 19 citations
- StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History SelectionSara Papi, Marco Gaido, Matteo Negri, Luisa BentivogliACL 2024
- HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationDavid Dale, Elena Voita, Janice Lam, Prangthip Hansanti et al.EMNLP 2023 · 8 citations
