Phone Features Improve Speech Translation
Elizabeth Salesky, Alan W. Black
Abstract
End-to-end models for speech translation (ST) more tightly couple speech recognition (ASR) and machine translation (MT) than a traditional cascade of separate ASR and MT models, with simpler model architectures and the potential for reduced error propagation. Their performance is often assumed to be superior, though in many conditions this is not yet the case. We compare cascaded and end-to-end models across high, medium, and low-resource conditions, and show that cascades remain stronger baselines. Further, we introduce two methods to incorporate phone features into ST models. We show that these features improve both architectures, closing the gap between end-to-end models and cascades, and outperforming previous academic work -by up to 9 BLEU on our low-resource setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Speech Translation and the End-to-End Promise: Taking Stock of Where We AreMatthias Sperber, Matthias PaulikACL 2020 · 7 citations
- Direct Segmentation Models for Streaming Speech TranslationJavier Iranzo-Sánchez, Adrià Giménez-Pastor, Joan Albert Silvestre-Cerdà, Pau Baquero-Arnal et al.EMNLP 2020 · 24 citations
- Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation EncodersChen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang et al.ACL 2021
- Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingYuchen Liu, Jiajun Zhang, Hao Xiong, Long Zhou et al.AAAI 2020 · 73 citations
- SpeechQE: Estimating the Quality of Direct Speech TranslationHyoJung Han, Kevin Duh, Marine CarpuatEMNLP 2024 · 1 citation
