A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
Yusheng Dai, Hang Chen, Jun Du, Ruoyu Wang, Shihao Chen, Haotian Wang, Chin-Hui Lee
摘要
Advanced Audio-Visual Speech Recognition (AVSR) systems have been observed to be sensitive to missing video frames, performing even worse than single-modality models. While applying the common dropout techniques to the video modality enhances robustness to missing frames, it simultaneously results in a performance loss when dealing with complete data input. In this study, we delve into this contrasting phenomenon through the lens of modality bias and uncover that an excessive modality bias towards the audio modality induced by dropout constitutes the fundamental cause. Next, we present the Modality Bias Hypothesis (MBH) to systematically describe the relationship between the modality bias and the robustness against missing modality in multimodal systems. Building on these findings, we propose a novel Multimodal Distribution Approximation with Knowledge Distillation (MDA-KD) framework to reduce over-reliance on the audio modality, maintaining performance and robustness simultaneously. Finally, to address an entirely missing modality, we adopt adapters to dynamically switch decision strategies. The effectiveness of our proposed approach is evaluated through comprehensive experiments on the MISP2021 and MISP2022 datasets. Our code is available at https://github . com/dalision/ModalBiasAVSR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual CuesTianrui Pan, Jie Liu, Bohan Wang, Jie Tang 等ACM MM 2024 · 被引用 3 次
- Hierarchical Codec Diffusion for Video-to-Speech GenerationJiaxin Ye, Gaoxiang Cong, Chenhui Wang, Xin-Cheng Wen 等CVPR 2026 · 被引用 3 次
- When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance CollapseYihuan Huang, Jun Xue, Liu Jiajun, Daixian Li 等CVPR 2026 · 被引用 2 次
- Adaptive Re-calibration Learning for Balanced Multimodal Intention RecognitionQu Yang, Xiyang Li, Fu Lin, Mang YeNeurIPS 2025 · 被引用 2 次
- MISP-Meeting: A Real-World Dataset with Multimodal Cues for Long-form Meeting Transcription and SummarizationHang Chen, Chao-Han Huck Yang, Jia-Chen Gu, Sabato Marco Siniscalchi 等ACL 2025
它引用的顶会 Paper16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou 等ICCV 2019 · 被引用 625 次
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov 等AAAI 2021 · 被引用 393 次
- Are Multimodal Transformers Robust to Missing Modality?Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine 等CVPR 2022 · 被引用 153 次
相关 Paper
- MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge DistillationHui Li, Pengfei Yang, Juanyang Chen, Le Dong 等ACM MM 2025 · 被引用 4 次
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang 等ICLR 2025
- Learnable Irrelevant Modality Dropout for Multimodal Action Recognition on Modality-Specific Annotated VideosSaghir Alfasly, Jian Lu, Chen Xu, Yuru ZouCVPR 2022 · 被引用 28 次
- Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete ModalitiesMingcheng Li, Dingkang Yang, Xiao Zhao, Shuaibing Wang 等CVPR 2024
- Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype DistillationJiaqi Tan, Xu Zheng, Yang LiuCVPR 2026 · 被引用 1 次
