MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition
Xize Cheng, Tao Jin, Rongjie Huang, Linjun Li, Wang Lin, Zehan Wang, Ye Wang, Huadai Liu, Aoxiong Yin, Zhou Zhao
Abstract
Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on visual speech. This lack of research is mainly due to the absence of datasets containing visual speech and translated text pairs. In this paper, we present AVMuST-TED, the first dataset for Audio-Visual Multilingual Speech Translation, derived from TED talks. Nonetheless, visual speech is not as distinguishable as audio speech, making it difficult to develop a mapping from source speech phonemes to the target language text. To address this issue, we propose MixSpeech, a cross-modality self-learning framework that utilizes audio speech to regularize the training of visual speech tasks. To further minimize the cross-modality gap and its impact on knowledge transfer, we suggest adopting mixed speech, which is created by interpolating audio and visual streams, along with a curriculum learning strategy to adjust the mixing ratio as needed. MixSpeech enhances speech translation in noisy environments, improving BLEU scores for four languages on AVMuST-TED by +1.4 to +4.2. Moreover, it achieves state-of-the-art performance in lip reading on CMLR (11.1%), LRS2 (25.5%), and LRS3 (28.0%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98cd0804-4fa2-409b-affb-29ac4d379266Cited by top-tier papers11
- Geodesic Multi-Modal Mixup for Robust Fine-TuningChangdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim et al.NeurIPS 2023 · 49 citations
- OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality AlignmentXize Cheng, Tao Jin, Linjun Li, Wang Lin et al.ACL 2023 · 10 citations
- SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local EditingLingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu et al.ACM MM 2024 · 10 citations
- XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech PerceptionHyoJung Han, Mohamed Anwar, Juan Pino, Wei-Ning Hsu et al.ACL 2024 · 9 citations
- TAVT: Towards Transferable Audio-Visual Text GenerationWang Lin, Tao Jin, Wenwen Pan, Linjun Li et al.ACL 2023 · 9 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
Related papers
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
- AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech RepresentationJeongsoo Choi, Se Jin Park, Minsu Kim, Yong Man RoCVPR 2024
- AV-TranSpeech: Audio-Visual Robust Speech-to-Speech TranslationRongjie Huang, Huadai Liu, Xize Cheng, Yi Ren et al.ACL 2023 · 9 citations
- CMOT: Cross-modal Mixup via Optimal Transport for Speech TranslationYan Zhou, Qingkai Fang, Yang FengACL 2023 · 24 citations
- Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationYuhao Zhang, Chen Xu, Bei Li, Hao Chen et al.EMNLP 2023 · 4 citations
