When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse
Yihuan Huang, Jun Xue, Liu Jiajun, Daixian Li, Tong Zhang, Zhuolin Yi, Yanzhen Ren, Kai Li
Abstract
Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored. This paper presents the first systematic evaluation of state-of-the-art AVSR models across mainstream VC platforms, revealing severe performance degradation caused by transmission distortions and spontaneous human hyperexpression. To address this gap, we construct MLD-VC, the first multimodal dataset tailored for VC, comprising 31 speakers, 22.79 hours of audio-visual data, and explicit use of the Lombard effect to enhance human hyper-expression. Through comprehensive analysis, we find that speech enhancement algorithms are the primary source of distribution shift, which alters the first and second formants of audio. Interestingly, we find that the distribution shift induced by the Lombard effect closely resembles that introduced by speech enhancement, which explains why models trained on Lombard data exhibit greater robustness in VC. Fine-tuning AVSR models on MLD-VC mitigates this issue, achieving an average 17.5% reduction in CER across several VC platforms. Our findings and dataset provide a foundation for developing more robust and generalizable AVSR systems in real-world video conferencing. MLD-VC is available at https://huggingface.co/datasets/nccm2p2/MLD-VC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4760ae77-1a10-4bf7-908e-eed125adb501Builds on7
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech SeparationUi-Hyeop Shin, Sangyoun Lee, Taehan Kim, Hyung-Min ParkNeurIPS 2024 · 46 citations
- Visual Hallucination Elevates Speech RecognitionFang Zhang, Yongxin Zhu, Xiangxiang Wang, Huang Chen et al.AAAI 2024 · 5 citations
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis et al.ICCV 2025 · 3 citations
- A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech RecognitionYusheng Dai, Hang Chen, Jun Du, Ruoyu Wang et al.CVPR 2024
Related papers
- Lombard-VLD: Voice Liveness Detection Based on Human Auditory FeedbackHongcheng Zhu, Zongkun Sun, Yanzhen Ren, Kun He et al.S&P 2025
- Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented PromptsDongjie Fu, Xize Cheng, Xiaoda Yang, Hanting Wang et al.ACM MM 2024 · 5 citations
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang et al.ICLR 2025
- Uni-DocRobust: Universal Plug-and-Play Robustness Enhancement for Multi-modal LLMs via Feature RestorationYuxuan Zhou, Baole Wei, Xingjian Hu, Haowei Chen et al.ICML 2026
- MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and RecognitionXize Cheng, Tao Jin, Rongjie Huang, Linjun Li et al.ICCV 2023 · 30 citations
