XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
HyoJung Han, Mohamed Anwar, Juan Pino, Wei-Ning Hsu, Marine Carpuat, Bowen Shi, Changhan Wang
Abstract
Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potential to improve robustness to noise. However, audio-visual (AV) data is only available in limited amounts and for fewer languages than audio-only resources. To address this gap, we present XLAVS-R, a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages. It is designed to maximize the benefits of limited multilingual AV pre-training data, by building on top of audio-only multilingual pre-training and simplifying existing pre-training schemes. Extensive evaluation on the MuAViC benchmark shows the strength of XLAVS-R on downstream audio-visual speech recognition and translation tasks, where it outperforms the previous state of the art by up to 18.5% WER and 4.7 BLEU given noisy AV inputs, and enables strong zero-shot audio-visual ability with audio-only fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48150bc4-dc3d-44fc-b636-00ce0f9bce27Cited by top-tier papers6
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech RecognitionUmberto Cappellazzo, Minsu Kim, Pingchuan Ma, Honglie Chen et al.NeurIPS 2025 · 5 citations
- AudioVSR: Enhancing Video Speech Recognition with Audio DataXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu et al.EMNLP 2024 · 3 citations
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis et al.ICCV 2025 · 3 citations
- SpeechQE: Estimating the Quality of Direct Speech TranslationHyoJung Han, Kevin Duh, Marine CarpuatEMNLP 2024 · 1 citation
- Measuring User's Mental Models of Speech Translation in Human-AI CollaborationHyojung Han, Nishant Balepur, Jordan Lee Boyd-Graber, Marine CarpuatACL 2026
Builds on8
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and RecognitionXize Cheng, Tao Jin, Rongjie Huang, Linjun Li et al.ICCV 2023 · 30 citations
- Jointly Learning Visual and Auditory Speech Representations from Raw DataAlexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis et al.ICLR 2023 · 13 citations
Related papers
- AV-TranSpeech: Audio-Visual Robust Speech-to-Speech TranslationRongjie Huang, Huadai Liu, Xize Cheng, Yi Ren et al.ACL 2023 · 9 citations
- AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech RepresentationJeongsoo Choi, Se Jin Park, Minsu Kim, Yong Man RoCVPR 2024
- Mu2SLAM: Multitask, Multilingual Speech and Language ModelsYong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey et al.ICML 2023 · 10 citations
- MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech RecognitionSungnyun Kim, Kangwook Jang, Sangmin Bae, Sungwoo Cho et al.ICML 2025
- Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech RecognitionYuchen Hu, Ruizhe Li, Chen Chen, Chengwei Qin et al.ACL 2023 · 7 citations
