Pre-training for Speech Translation: CTC Meets Optimal Transport
Phuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino, Benjamin Lecouteux, Didier Schwab
Abstract
The gap between speech and text modalities is a major challenge in speech-to-text translation (ST). Different methods have been proposed to reduce this gap, but most of them require architectural changes in ST training. In this work, we propose to mitigate this issue at the pre-training stage, requiring no change in the ST model. First, we show that the connectionist temporal classification (CTC) loss can reduce the modality gap by design. We provide a quantitative comparison with the more common cross-entropy loss, showing that pre-training with CTC consistently achieves better final ST accuracy. Nevertheless, CTC is only a partial solution and thus, in our second contribution, we propose a novel pre-training method combining CTC and optimal transport to further reduce this gap. Our method pre-trains a Siamese-like model composed of two encoders, one for acoustic inputs and the other for textual inputs, such that they produce representations that are close to each other in the Wasserstein space. Extensive experiments on the standard CoVoST-2 and MuST-C datasets show that our pre-training method applied to the vanilla encoder-decoder Transformer achieves state-of-the-art performance under the no-external-data setting, and performs on par with recent strong multi-task learning systems trained with external data. Finally, our method can also be applied on top of these multi-task systems, leading to further improvements for these models. Code and pre-trained models are available at https://github.com/formiel/fairseq.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 404b9571-e73c-4ef0-91a2-ea04e634a01cCited by top-tier papers6
- Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationYuhao Zhang, Chen Xu, Bei Li, Hao Chen et al.EMNLP 2023 · 4 citations
- Improving Language and Modality Transfer in Translation by Character-level ModelingIoannis Tsiamas, David Dale, Marta R. Costa-jussàACL 2025 · 3 citations
- PromptST: Abstract Prompt Learning for End-to-End Speech TranslationTengfei Yu, Liang Ding, Xuebo Liu, Kehai Chen et al.EMNLP 2023 · 3 citations
- PLaST: Towards Paralinguistic-aware Speech TranslationYi Li, Rui Zhao, Ruiquan Zhang, Jinsong Su et al.AAAI 2026
- Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum LearningYexing Du, Youcheng Pan, Ziyang Ma, Bo Yang et al.ACL 2025
Builds on11
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou et al.ACL 2020 · 100 citations
- Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang et al.AAAI 2020 · 90 citations
- Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text TranslationQianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou et al.AAAI 2021 · 65 citations
- Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal TransportHsin-Yi Lin, Huan-Hsin Tseng, Xugang Lu, Yu TsaoNeurIPS 2021 · 40 citations
Related papers
- CMOT: Cross-modal Mixup via Optimal Transport for Speech TranslationYan Zhou, Qingkai Fang, Yang FengACL 2023 · 24 citations
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang et al.AAAI 2023 · 17 citations
- Understanding and Bridging the Modality Gap for Speech TranslationQingkai Fang, Yang FengACL 2023 · 12 citations
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu et al.EMNLP 2022 · 38 citations
- Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation TaskYun Tang, Juan Miguel Pino, Xian Li, Changhan Wang et al.ACL 2021
