Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical Errors
Nikita Mehandru, Sweta Agrawal, Yimin Xiao, Ge Gao, Elaine C. Khoong, Marine Carpuat, Niloufar Salehi
摘要
A major challenge in the practical use of Machine Translation (MT) is that users lack information on translation quality to make informed decisions about how to rely on outputs. Progress in quality estimation research provides techniques to automatically assess MT quality, but these techniques have primarily been evaluated in vitro by comparison against human judgments outside of a specific context of use. This paper evaluates quality estimation feedback in vivo with a human study in realistic high-stakes medical settings. Using Emergency Department discharge instructions, we study how interventions based on quality estimation versus backtranslation assist physicians in deciding whether to show MT outputs to a patient. We find that quality estimation improves appropriate reliance on MT, but backtranslation helps physicians detect more clinically harmful errors that QE alone often misses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Sustaining Human Agency, Attending to Its Cost: An Investigation into Generative AI Design for Non-Native Speakers' Language UseYimin Xiao, Cartor Hancock, Sweta Agrawal, Nikita Mehandru 等CHI 2025 · 被引用 18 次
- The News Says, the Bot Says: How Immigrants and Locals Differ in Chatbot-Facilitated News ReadingYongle Zhang, Phuong-Anh Nguyen-Le, Kriti Singh, Ge GaoCHI 2025 · 被引用 5 次
- An Interdisciplinary Approach to Human-Centered Machine TranslationMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli 等EMNLP 2025 · 被引用 2 次
- Designing Beyond Language: Sociotechnical Barriers in AI Health Technologies for Limited English ProficiencyMichelle Huang, Violeta J. Rodriguez, Koustuv Saha, Tal AugustCHI 2026 · 被引用 2 次
- SpeechQE: Estimating the Quality of Direct Speech TranslationHyoJung Han, Kevin Duh, Marine CarpuatEMNLP 2024 · 被引用 1 次
它引用的顶会 Paper5
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 被引用 962 次
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok 等CHI 2021 · 被引用 713 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 被引用 47 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
相关 Paper
- Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine TranslationDayeon Ki, Kevin Duh, Marine CarpuatEMNLP 2025
- Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation OutputRaksha Shenoy, Nico Herbig, Antonio Krüger, Josef van GenabithEMNLP 2021 · 被引用 3 次
- Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect TranslationsYimin Xiao, Yongle Zhang, Dayeon Ki, Calvin Bao 等EMNLP 2025
- Viability of Machine Translation for Healthcare in Low-Resourced LanguagesHellina Hailu Nigatu, Nikita Mehandru, Negasi Haile Abadi, Blen Gebremeskel 等EMNLP 2025
- Measuring User's Mental Models of Speech Translation in Human-AI CollaborationHyojung Han, Nishant Balepur, Jordan Lee Boyd-Graber, Marine CarpuatACL 2026
