Toward Machine Interpreting: Lessons from Human Interpreting Studies
Matthias Sperber, Maureen de Seyssel, Jiajun Bao, Matthias Paulik
Abstract
Current speech translation systems, while having achieved impressive accuracies, are rather static in their behavior and do not adapt to realworld situations in ways human interpreters do. In order to improve their practical usefulness and enable interpreting-like experiences, a precise understanding of the nature of human interpreting is crucial. To this end, we discuss human interpreting literature from the perspective of the machine translation field, while considering both operational and qualitative aspects. We identify implications for the development of speech translation systems and argue that there is great potential to adopt many human interpreting principles using recent modeling techniques. We hope that our findings provide inspiration for closing the perceived usability gap, and can motivate progress toward true machine interpreting. Feature Description Example Temporal immediacy Produces interpretation in real-time. Maintains 1-2 seconds ear-voice-span. Spatial immediacy Operates in proximity of speaker & audience. Shares stage with speaker. Multimodality Uses visual or gesture cues when available. Refers to chart while speaker points. Free/diverse actions Dynamically adapts to any situation. Adapts translation approach to content type. Interaction/influence Acts as a independent agent when needed. Requests clarification, improves acoustics. Intent translation Interprets what is meant, not what is said. Interpretation conveys hidden accusations. Interpreter uncertainty Maintains trust by signaling own uncertainty. "Speaker may have said 'revenue'." Speaker errors Indicates or corrects unintentional speaker errors. Corrects "million" to "billion" in context. Adaptation/explanation Adapts or explains culture-specific expressions. "Break a leg!" → "Good luck!". Explicitation Explicitates logic, intent, order, viewpoints. "..., according to X's view." Brevity Keeps sentences short and clear. "The results were strong. More tests needed." Rhetoric quality Delivers exceptionally high rhetoric quality. Adapts style to particular audience. Pleasant experience Works reliably; pleasant voice; eye contact. Avoids hectic speech when falling behind. Cognitive ergonomics Minimizes audience stress and fatigue. Avoids complex language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc19f2a2-aed3-4c76-8727-3fe12fb0e2baBuilds on17
- CultureLLM: Incorporating Cultural Differences into Large Language ModelsCheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram et al.NeurIPS 2024 · 101 citations
- Language Model Can Listen While SpeakingZiyang Ma, Yakun Song, Chenpeng Du, Jian Cong et al.AAAI 2025 · 58 citations
- Unmet Needs and Opportunities for Mobile Translation AIDaniel J. Liebling, Michal Lahav, Abigail Evans, Aaron Donsbach et al.CHI 2020 · 47 citations
- Automatic Fact-Guided Sentence ModificationDarsh J. Shah, Tal Schuster, Regina BarzilayAAAI 2020 · 44 citations
- Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge GraphsSimone Conia, Daniel Lee, Min Li, Umar Farooq Minhas et al.EMNLP 2024 · 7 citations
Related papers
- An Interdisciplinary Approach to Human-Centered Machine TranslationMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli et al.EMNLP 2025 · 2 citations
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 26 citations
- SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question AnsweringHyoJung Han, Marine Carpuat, Jordan L. Boyd-GraberEMNLP 2022 · 3 citations
- Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect TranslationsYimin Xiao, Yongle Zhang, Dayeon Ki, Calvin Bao et al.EMNLP 2025
- High-Fidelity Simultaneous Speech-To-Speech TranslationTom Labiausse, Laurent Mazaré, Edouard Grave, Alexandre Défossez et al.ICML 2025
