Speech Translation and the End-to-End Promise: Taking Stock of Where We Are
Matthias Sperber, Matthias Paulik
摘要
Over its three decade history, speech translation has experienced several shifts in its primary research themes; moving from loosely coupled cascades of speech recognition and machine translation, to exploring questions of tight coupling, and finally to end-to-end models that have recently attracted much attention. This paper provides a brief survey of these developments, along with a discussion of the main challenges of traditional approaches which stem from committing to intermediate representations from the speech recognizer, and from training cascaded models separately towards different objectives. Recent end-to-end modeling techniques promise a principled way of overcoming these issues by allowing joint training of all model components and removing the need for explicit intermediate representations. However, a closer look reveals that many end-to-end models fall short of solving these issues, due to compromises made to address data scarcity. This paper provides a unifying categorization and nomenclature that covers both traditional and recent approaches and that may help researchers by highlighting both trade-offs and open research questions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Consecutive Decoding for Speech-to-text TranslationQianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu 等AAAI 2021 · 被引用 46 次
- TransVIP: Speech to Speech Translation System with Voice and Isochrony PreservationChenyang Le, Yao Qian, Dongmei Wang, Long Zhou 等NeurIPS 2024 · 被引用 25 次
- ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text TranslationChenyang Le, Yao Qian, Long Zhou, Shujie Liu 等NeurIPS 2023 · 被引用 21 次
- Understanding and Bridging the Modality Gap for Speech TranslationQingkai Fang, Yang FengACL 2023 · 被引用 12 次
- Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?Marco Gaido, Sara Papi, Matteo Negri, Luisa BentivogliACL 2024 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Phone Features Improve Speech TranslationElizabeth Salesky, Alan W. BlackACL 2020 · 被引用 2 次
- Summarizing Speech: A Comprehensive SurveyFabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau 等EMNLP 2025 · 被引用 3 次
- In-Situ Text-Only Adaptation of Speech Models with Low-Overhead Speech ImputationsAshish R. Mittal, Sunita Sarawagi, Preethi JyothiICLR 2023
- SpeechQE: Estimating the Quality of Direct Speech TranslationHyoJung Han, Kevin Duh, Marine CarpuatEMNLP 2024 · 被引用 1 次
- Recent Advances in Speech Language Models: A SurveyWenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng 等ACL 2025
