Unified Speech-Text Pre-training for Speech Translation and Recognition
Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, Juan Miguel Pino
摘要
In this work, we describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method utilizes multi-task learning to integrate four self-supervised and supervised subtasks for cross modality learning. A self-supervised speech subtask, which leverages unlabelled speech data, and a (self-)supervised text to text subtask, which makes use of abundant text training data, take up the majority of the pre-training time. Two auxiliary supervised speech tasks are included to unify speech and text modeling space. Detailed analysis reveals learning interference among subtasks. In order to alleviate the subtask interference, two pre-training configurations are proposed for speech translation and speech recognition respectively. Our experiments show the proposed method can effectively fuse speech and text information into one model. It achieves between 1.7 and 2.3 BLEU improvement above the state of the art on the MuST-C speech translation dataset and comparable WERs to wav2vec 2.0 on the Librispeech speech recognition task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu 等EMNLP 2022 · 被引用 38 次
- XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech CodecsYitian Gong, Luozhijie Jin, Kuangwei Chen, Dong Zhang 等ACL 2026 · 被引用 35 次
- Pre-training for Speech Translation: CTC Meets Optimal TransportPhuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino 等ICML 2023 · 被引用 33 次
- UnitY: Two-pass Direct Speech-to-speech Translation with Discrete UnitsHirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen 等ACL 2023 · 被引用 30 次
- Speech-Text Pre-training for Spoken Dialog Understanding with Explicit Cross-Modal AlignmentTianshu Yu, Haoyu Gao, Ting-En Lin, Min Yang 等ACL 2023 · 被引用 26 次
它引用的顶会 Paper4
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 被引用 74 次
- Multilingual Speech Translation from Efficient Finetuning of Pretrained ModelsXian Li, Changhan Wang, Yun Tang, Chau Tran 等ACL 2021
相关 Paper
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang 等AAAI 2023 · 被引用 17 次
- Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation TaskYun Tang, Juan Miguel Pino, Xian Li, Changhan Wang 等ACL 2021
- T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine TranslationPaul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger SchwenkEMNLP 2022 · 被引用 8 次
- Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationYuhao Zhang, Chen Xu, Bei Li, Hao Chen 等EMNLP 2023 · 被引用 4 次
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov 等ACL 2023 · 被引用 9 次
