Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical Errors
Nikita Mehandru, Sweta Agrawal, Yimin Xiao, Ge Gao, Elaine C. Khoong, Marine Carpuat, Niloufar Salehi
Abstract
A major challenge in the practical use of Machine Translation (MT) is that users lack information on translation quality to make informed decisions about how to rely on outputs. Progress in quality estimation research provides techniques to automatically assess MT quality, but these techniques have primarily been evaluated in vitro by comparison against human judgments outside of a specific context of use. This paper evaluates quality estimation feedback in vivo with a human study in realistic high-stakes medical settings. Using Emergency Department discharge instructions, we study how interventions based on quality estimation versus backtranslation assist physicians in deciding whether to show MT outputs to a patient. We find that quality estimation improves appropriate reliance on MT, but backtranslation helps physicians detect more clinically harmful errors that QE alone often misses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b763844e-74c7-48e4-9e6e-4bec95e6daccCited by top-tier papers10
- Sustaining Human Agency, Attending to Its Cost: An Investigation into Generative AI Design for Non-Native Speakers' Language UseYimin Xiao, Cartor Hancock, Sweta Agrawal, Nikita Mehandru et al.CHI 2025 · 18 citations
- The News Says, the Bot Says: How Immigrants and Locals Differ in Chatbot-Facilitated News ReadingYongle Zhang, Phuong-Anh Nguyen-Le, Kriti Singh, Ge GaoCHI 2025 · 5 citations
- An Interdisciplinary Approach to Human-Centered Machine TranslationMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli et al.EMNLP 2025 · 2 citations
- Designing Beyond Language: Sociotechnical Barriers in AI Health Technologies for Limited English ProficiencyMichelle Huang, Violeta J. Rodriguez, Koustuv Saha, Tal AugustCHI 2026 · 2 citations
- SpeechQE: Estimating the Quality of Direct Speech TranslationHyoJung Han, Kevin Duh, Marine CarpuatEMNLP 2024 · 1 citation
Builds on5
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 962 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 47 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine TranslationDayeon Ki, Kevin Duh, Marine CarpuatEMNLP 2025
- Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation OutputRaksha Shenoy, Nico Herbig, Antonio Krüger, Josef van GenabithEMNLP 2021 · 3 citations
- Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect TranslationsYimin Xiao, Yongle Zhang, Dayeon Ki, Calvin Bao et al.EMNLP 2025
- Viability of Machine Translation for Healthcare in Low-Resourced LanguagesHellina Hailu Nigatu, Nikita Mehandru, Negasi Haile Abadi, Blen Gebremeskel et al.EMNLP 2025
- Measuring User's Mental Models of Speech Translation in Human-AI CollaborationHyojung Han, Nishant Balepur, Jordan Lee Boyd-Graber, Marine CarpuatACL 2026
