Consecutive Decoding for Speech-to-text Translation
Qianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu, Bo Xu, Lei Li
摘要
Speech-to-text translation (ST), which directly translates the source language speech to the target language text, has attracted intensive attention recently. However, the combination of speech recognition and machine translation in a single model poses a heavy burden on the direct cross-modal crosslingual mapping. To reduce the learning difficulty, we propose COnSecutive Transcription and Translation (COSTT), an integral approach for speech-to-text translation. The key idea is to generate source transcript and target translation text with a single decoder. It benefits the model training so that additional large parallel text corpus can be fully exploited to enhance the speech translation training. Our method is verified on three mainstream datasets, including Augmented LibriSpeech English-French dataset, IWSLT2018 English-German dataset, and TED English-Chinese dataset. Experiments show that our proposed COSTT outperforms or on par with the previous state-of-the-art methods on the three datasets. We have released our code at https://github.com/ dqqcasia/st .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 被引用 74 次
- PEIT: Bridging the Modality Gap with Pre-trained Models for End-to-End Image TranslationShaolin Zhu, Shangjie Li, Yikun Lei, Deyi XiongACL 2023 · 被引用 12 次
- Non-Parametric Domain Adaptation for End-to-End Speech TranslationYichao Du, Weizhi Wang, Zhirui Zhang, Boxing Chen 等EMNLP 2022 · 被引用 10 次
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang 等AAAI 2024 · 被引用 6 次
- Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyLinlin Zhang, Kai Fan, Jiajun Bu, Zhongqiang HuangEMNLP 2023 · 被引用 1 次
它引用的顶会 Paper3
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou 等ACL 2020 · 被引用 100 次
- Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang 等AAAI 2020 · 被引用 90 次
- Speech Translation and the End-to-End Promise: Taking Stock of Where We AreMatthias Sperber, Matthias PaulikACL 2020 · 被引用 7 次
相关 Paper
- Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingYuchen Liu, Jiajun Zhang, Hao Xiong, Long Zhou 等AAAI 2020 · 被引用 73 次
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang 等ACL 2022 · 被引用 104 次
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu 等EMNLP 2022 · 被引用 38 次
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov 等ACL 2023 · 被引用 9 次
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang 等AAAI 2023 · 被引用 17 次
