When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse
Yihuan Huang, Jun Xue, Liu Jiajun, Daixian Li, Tong Zhang, Zhuolin Yi, Yanzhen Ren, Kai Li
摘要
Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored. This paper presents the first systematic evaluation of state-of-the-art AVSR models across mainstream VC platforms, revealing severe performance degradation caused by transmission distortions and spontaneous human hyperexpression. To address this gap, we construct MLD-VC, the first multimodal dataset tailored for VC, comprising 31 speakers, 22.79 hours of audio-visual data, and explicit use of the Lombard effect to enhance human hyper-expression. Through comprehensive analysis, we find that speech enhancement algorithms are the primary source of distribution shift, which alters the first and second formants of audio. Interestingly, we find that the distribution shift induced by the Lombard effect closely resembles that introduced by speech enhancement, which explains why models trained on Lombard data exhibit greater robustness in VC. Fine-tuning AVSR models on MLD-VC mitigates this issue, achieving an average 17.5% reduction in CER across several VC platforms. Our findings and dataset provide a foundation for developing more robust and generalizable AVSR systems in real-world video conferencing. MLD-VC is available at https://huggingface.co/datasets/nccm2p2/MLD-VC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech SeparationUi-Hyeop Shin, Sangyoun Lee, Taehan Kim, Hyung-Min ParkNeurIPS 2024 · 被引用 46 次
- Visual Hallucination Elevates Speech RecognitionFang Zhang, Yongxin Zhu, Xiangxiang Wang, Huang Chen 等AAAI 2024 · 被引用 5 次
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis 等ICCV 2025 · 被引用 3 次
- A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech RecognitionYusheng Dai, Hang Chen, Jun Du, Ruoyu Wang 等CVPR 2024
相关 Paper
- Lombard-VLD: Voice Liveness Detection Based on Human Auditory FeedbackHongcheng Zhu, Zongkun Sun, Yanzhen Ren, Kun He 等S&P 2025
- Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented PromptsDongjie Fu, Xize Cheng, Xiaoda Yang, Hanting Wang 等ACM MM 2024 · 被引用 5 次
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang 等ICLR 2025
- Uni-DocRobust: Universal Plug-and-Play Robustness Enhancement for Multi-modal LLMs via Feature RestorationYuxuan Zhou, Baole Wei, Xingjian Hu, Haowei Chen 等ICML 2026
- MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and RecognitionXize Cheng, Tao Jin, Rongjie Huang, Linjun Li 等ICCV 2023 · 被引用 30 次
