STEMM: Self-learning with Speech-text Manifold Mixup for Speech Translation
Qingkai Fang, Rong Ye, Lei Li, Yang Feng, Mingxuan Wang
摘要
How to learn a better speech representation for end-to-end speech-to-text translation (ST) with limited labeled data? Existing techniques often attempt to transfer powerful machine translation (MT) capabilities to ST, but neglect the representation discrepancy across modalities. In this paper, we propose the Speech-TExt Manifold Mixup (STEMM) method to calibrate such discrepancy. Specifically, we mix up the representation sequences of different modalities, and take both unimodal speech sequences and multimodal mixed sequences as input to the translation model in parallel, and regularize their output predictions with a selflearning framework. Experiments on MuST-C speech translation benchmark and further analysis show that our method effectively alleviates the cross-modal representation discrepancy, and achieves significant improvements over a strong baseline on eight translation directions. * indicates corresponding authors. † Work was done while at ByteDance AI Lab. Part of joint project between ICT/CAS and ByteDance AI Lab. Work was done when QF was a member of the joint project. Code and models are publicly available at https:// github.com/ictnlp/STEMM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Geodesic Multi-Modal Mixup for Robust Fine-TuningChangdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim 等NeurIPS 2023 · 被引用 49 次
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu 等EMNLP 2022 · 被引用 38 次
- Pre-training for Speech Translation: CTC Meets Optimal TransportPhuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino 等ICML 2023 · 被引用 33 次
- CMOT: Cross-modal Mixup via Optimal Transport for Speech TranslationYan Zhou, Qingkai Fang, Yang FengACL 2023 · 被引用 24 次
- DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech TranslationQingkai Fang, Yan Zhou, Yang FengNeurIPS 2023 · 被引用 22 次
它引用的顶会 Paper11
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou 等ACL 2020 · 被引用 100 次
- Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang 等AAAI 2020 · 被引用 90 次
相关 Paper
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 被引用 74 次
- Understanding and Bridging the Modality Gap for Speech TranslationQingkai Fang, Yang FengACL 2023 · 被引用 12 次
- Regularizing End-to-End Speech Translation with Triangular Decomposition AgreementYichao Du, Zhirui Zhang, Weizhi Wang, Boxing Chen 等AAAI 2022 · 被引用 25 次
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang 等ACL 2022 · 被引用 104 次
- Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationYuhao Zhang, Chen Xu, Bei Li, Hao Chen 等EMNLP 2023 · 被引用 4 次
