Understanding and Bridging the Modality Gap for Speech Translation
Qingkai Fang, Yang Feng
Abstract
How to achieve better end-to-end speech translation (ST) by leveraging (text) machine translation (MT) data? Among various existing techniques, multi-task learning is one of the effective ways to share knowledge between ST and MT in which additional MT data can help to learn source-to-target mapping. However, due to the differences between speech and text, there is always a gap between ST and MT. In this paper, we first aim to understand this modality gap from the target-side representation differences, and link the modality gap to another well-known problem in neural machine translation: exposure bias. We find that the modality gap is relatively small during training except for some difficult cases, but keeps increasing during inference due to the cascading effect. To address these problems, we propose the Cross-modal Regularization with Scheduled Sampling (Cress) method. Specifically, we regularize the output predictions of ST and MT, whose target-side contexts are derived by sampling between ground truth words and self-generated words with a varying probability. Furthermore, we introduce token-level adaptive training which assigns different training weights to target tokens to handle difficult cases with large modality gaps. Experiments and analysis show that our approach effectively bridges the modality gap, and achieves significant improvements over a strong baseline in all eight directions of the MuST-C dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech TranslationQingkai Fang, Yan Zhou, Yang FengNeurIPS 2023 · 22 citations
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 9 citations
- Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language ModelsBajian Xiang, Shuaijiang Zhao, Tingwei Guo, Wei ZouEMNLP 2025 · 6 citations
- Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationYuhao Zhang, Chen Xu, Bei Li, Hao Chen et al.EMNLP 2023 · 4 citations
- Curriculum Consistency Learning for Conditional Sentence GenerationLiangxin Liu, Xuebo Liu, Lian Lian, Shengjun Cheng et al.EMNLP 2024 · 1 citation
Builds on19
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- R-Drop: Regularized Dropout for Neural NetworksXiaobo Liang, Lijun Wu, Juntao Li, Yue Wang et al.NeurIPS 2021 · 610 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou et al.ACL 2020 · 100 citations
Related papers
- CMOT: Cross-modal Mixup via Optimal Transport for Speech TranslationYan Zhou, Qingkai Fang, Yang FengACL 2023 · 24 citations
- Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation TaskYun Tang, Juan Miguel Pino, Xian Li, Changhan Wang et al.ACL 2021
- Regularizing End-to-End Speech Translation with Triangular Decomposition AgreementYichao Du, Zhirui Zhang, Weizhi Wang, Boxing Chen et al.AAAI 2022 · 25 citations
- Pre-training for Speech Translation: CTC Meets Optimal TransportPhuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino et al.ICML 2023 · 33 citations
- STEMM: Self-learning with Speech-text Manifold Mixup for Speech TranslationQingkai Fang, Rong Ye, Lei Li, Yang Feng et al.ACL 2022
