Toward Machine Interpreting: Lessons from Human Interpreting Studies
Matthias Sperber, Maureen de Seyssel, Jiajun Bao, Matthias Paulik
摘要
Current speech translation systems, while having achieved impressive accuracies, are rather static in their behavior and do not adapt to realworld situations in ways human interpreters do. In order to improve their practical usefulness and enable interpreting-like experiences, a precise understanding of the nature of human interpreting is crucial. To this end, we discuss human interpreting literature from the perspective of the machine translation field, while considering both operational and qualitative aspects. We identify implications for the development of speech translation systems and argue that there is great potential to adopt many human interpreting principles using recent modeling techniques. We hope that our findings provide inspiration for closing the perceived usability gap, and can motivate progress toward true machine interpreting. Feature Description Example Temporal immediacy Produces interpretation in real-time. Maintains 1-2 seconds ear-voice-span. Spatial immediacy Operates in proximity of speaker & audience. Shares stage with speaker. Multimodality Uses visual or gesture cues when available. Refers to chart while speaker points. Free/diverse actions Dynamically adapts to any situation. Adapts translation approach to content type. Interaction/influence Acts as a independent agent when needed. Requests clarification, improves acoustics. Intent translation Interprets what is meant, not what is said. Interpretation conveys hidden accusations. Interpreter uncertainty Maintains trust by signaling own uncertainty. "Speaker may have said 'revenue'." Speaker errors Indicates or corrects unintentional speaker errors. Corrects "million" to "billion" in context. Adaptation/explanation Adapts or explains culture-specific expressions. "Break a leg!" → "Good luck!". Explicitation Explicitates logic, intent, order, viewpoints. "..., according to X's view." Brevity Keeps sentences short and clear. "The results were strong. More tests needed." Rhetoric quality Delivers exceptionally high rhetoric quality. Adapts style to particular audience. Pleasant experience Works reliably; pleasant voice; eye contact. Avoids hectic speech when falling behind. Cognitive ergonomics Minimizes audience stress and fatigue. Avoids complex language.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- CultureLLM: Incorporating Cultural Differences into Large Language ModelsCheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram 等NeurIPS 2024 · 被引用 101 次
- Language Model Can Listen While SpeakingZiyang Ma, Yakun Song, Chenpeng Du, Jian Cong 等AAAI 2025 · 被引用 58 次
- Unmet Needs and Opportunities for Mobile Translation AIDaniel J. Liebling, Michal Lahav, Abigail Evans, Aaron Donsbach 等CHI 2020 · 被引用 47 次
- Automatic Fact-Guided Sentence ModificationDarsh J. Shah, Tal Schuster, Regina BarzilayAAAI 2020 · 被引用 44 次
- Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge GraphsSimone Conia, Daniel Lee, Min Li, Umar Farooq Minhas 等EMNLP 2024 · 被引用 7 次
相关 Paper
- An Interdisciplinary Approach to Human-Centered Machine TranslationMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli 等EMNLP 2025 · 被引用 2 次
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 被引用 26 次
- SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question AnsweringHyoJung Han, Marine Carpuat, Jordan L. Boyd-GraberEMNLP 2022 · 被引用 3 次
- Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect TranslationsYimin Xiao, Yongle Zhang, Dayeon Ki, Calvin Bao 等EMNLP 2025
- High-Fidelity Simultaneous Speech-To-Speech TranslationTom Labiausse, Laurent Mazaré, Edouard Grave, Alexandre Défossez 等ICML 2025
