AV-RISE: Hierarchical Cross-Modal Denoising for Learning Robust Audio-Visual Speech Representation
Zhishuo Zhao, Yi Lin, Dongyue Guo, Junyu Fan
Abstract
Audio-visual speech recognition (AVSR) leverages complementary visual cues to improve speech recognition. However, in real-world scenarios, both modalities may suffer from noise or occlusion. In such scenarios, most existing fusion strategies overlook the variation in modality-specific quality under different degradation conditions. This limitation may lead to dominance of corrupted modality in the fusion process, resulting in worse AVSR performance than unimodal systems, termed as Corrupted Modality Bias (CMB) in this work. To address this, a self-supervised speech representation learning framework, called AV-RISE, is proposed to employ teacher-student self-distillation to robustly reconstruct clean speech representations from corrupted audio-visual inputs. A hierarchical fusion mechanism is designed to progressively refine audio and visual representations by integrating the Suppression and Enhancement Interaction (SEI) module into each layer of the pre-trained encoder. In the SEI module, cross-modal suppression and modality-oriented enhancement are performed to mitigate noise-induced feature inconsistencies, which strengthens the modeling of complementary semantic representations. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that AV-RISE outperforms SOTA AVSR models, especially under extreme degradation. Most importantly, the hierarchical SEI-based fusion effectively enhances reliable semantic representations to mitigate CMB, by evaluating feature similarities between clean and noise samples.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b24ffe85-13fb-4214-83a0-3957a3bf4f08Related papers
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang et al.ICLR 2025
- Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement LearningChen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou et al.AAAI 2023 · 35 citations
- Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringJoanna Hong, Minsu Kim, Jeongsoo Choi, Yong Man RoCVPR 2023
- Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech RecognitionAlexandros Haliassos, Rodrigo Mira, Stavros PetridisICLR 2026 · 1 citation
- Visual Sound Localization in the Wild by Cross-Modal Interference ErasingXian Liu, Rui Qian, Hang Zhou, Di Hu et al.AAAI 2022 · 31 citations
