Measuring User's Mental Models of Speech Translation in Human-AI Collaboration
Hyojung Han, Nishant Balepur, Jordan Lee Boyd-Graber, Marine Carpuat
Abstract
Millions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do. This paper studies users' mental models of speech translation systems through a new framework based on cross-lingual question answering, where users either accept MT output or request professional re-translation to answer questions based on the information presented in a foreign language. By analyzing user behavior and accuracy trends across varying translation qualities, we examine to what extent they can predict where the system is likely to be wrong, and how this mental model evolves. Users develop stronger mental models with practice, especially when they have some knowledge of the source language, primarily by relying on surface-level error cues. Moreover, providing speech transcriptions can help users develop better mental models. Our results show the promise of cross-lingual question answering as a downstream task for studying MT mental models, and advancing our understanding of human-AI collaboration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab1dbf1d-7409-4f2d-b522-1c602c30542eBuilds on14
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic ScoresMaria De-Arteaga, Riccardo Fogliato, Alexandra ChouldechovaCHI 2020 · 176 citations
- Unmet Needs and Opportunities for Mobile Translation AIDaniel J. Liebling, Michal Lahav, Abigail Evans, Aaron Donsbach et al.CHI 2020 · 47 citations
- Sustaining Human Agency, Attending to Its Cost: An Investigation into Generative AI Design for Non-Native Speakers' Language UseYimin Xiao, Cartor Hancock, Sweta Agrawal, Nikita Mehandru et al.CHI 2025 · 18 citations
- Extrinsic Evaluation of Machine Translation MetricsNikita Moghe, Tom Sherborne, Mark Steedman, Alexandra BirchACL 2023 · 12 citations
Related papers
- Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation OutputRaksha Shenoy, Nico Herbig, Antonio Krüger, Josef van GenabithEMNLP 2021 · 3 citations
- Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect TranslationsYimin Xiao, Yongle Zhang, Dayeon Ki, Calvin Bao et al.EMNLP 2025
- Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical ErrorsNikita Mehandru, Sweta Agrawal, Yimin Xiao, Ge Gao et al.EMNLP 2023 · 9 citations
- Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine TranslationDayeon Ki, Kevin Duh, Marine CarpuatEMNLP 2025
- MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine TranslationParker Riley, Daniel Deutsch, Mara Finkelstein, Colten DiIanni et al.ACL 2026
