Consecutive Decoding for Speech-to-text Translation
Qianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu, Bo Xu, Lei Li
Abstract
Speech-to-text translation (ST), which directly translates the source language speech to the target language text, has attracted intensive attention recently. However, the combination of speech recognition and machine translation in a single model poses a heavy burden on the direct cross-modal crosslingual mapping. To reduce the learning difficulty, we propose COnSecutive Transcription and Translation (COSTT), an integral approach for speech-to-text translation. The key idea is to generate source transcript and target translation text with a single decoder. It benefits the model training so that additional large parallel text corpus can be fully exploited to enhance the speech translation training. Our method is verified on three mainstream datasets, including Augmented LibriSpeech English-French dataset, IWSLT2018 English-German dataset, and TED English-Chinese dataset. Experiments show that our proposed COSTT outperforms or on par with the previous state-of-the-art methods on the three datasets. We have released our code at https://github.com/ dqqcasia/st .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e32ee6ec-42b9-49e5-8d9d-a2ac3fd259a9Cited by top-tier papers7
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 74 citations
- PEIT: Bridging the Modality Gap with Pre-trained Models for End-to-End Image TranslationShaolin Zhu, Shangjie Li, Yikun Lei, Deyi XiongACL 2023 · 12 citations
- Non-Parametric Domain Adaptation for End-to-End Speech TranslationYichao Du, Weizhi Wang, Zhirui Zhang, Boxing Chen et al.EMNLP 2022 · 10 citations
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang et al.AAAI 2024 · 6 citations
- Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyLinlin Zhang, Kai Fan, Jiajun Bu, Zhongqiang HuangEMNLP 2023 · 1 citation
Builds on3
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou et al.ACL 2020 · 100 citations
- Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang et al.AAAI 2020 · 90 citations
- Speech Translation and the End-to-End Promise: Taking Stock of Where We AreMatthias Sperber, Matthias PaulikACL 2020 · 7 citations
Related papers
- Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingYuchen Liu, Jiajun Zhang, Hao Xiong, Long Zhou et al.AAAI 2020 · 73 citations
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu et al.EMNLP 2022 · 38 citations
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov et al.ACL 2023 · 9 citations
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang et al.AAAI 2023 · 17 citations
