Visual Hallucination Elevates Speech Recognition
Fang Zhang, Yongxin Zhu, Xiangxiang Wang, Huang Chen, Xing Sun, Linli Xu
Abstract
Due to the detrimental impact of noise on the conventional audio speech recognition (ASR) task, audio-visual speech recognition (AVSR) has been proposed by incorporating both audio and visual video signals. Although existing methods have demonstrated that the aligned visual input of lip movements can enhance the robustness of AVSR systems against noise, the paired videos are not always available during inference, leading to the problem of the missing visual modality, which restricts their practicality in real-world scenarios.
To tackle this problem, we propose a Discrete Feature based Visual Generative Model (DFVGM) which exploits semantic correspondences between the audio and visual modalities during training, generating visual hallucinations in lieu of real videos during inference. To achieve that, the primary challenge is to generate the visual hallucination given the noisy audio while preserving semantic correspondences with the clean speech. To tackle this challenge, we start with training the audio encoder in the Audio-Only (AO) setting, which generates continuous semantic features closely associated with the linguistic information. Simultaneously, the visual encoder is trained in the Visual-Only (VO) setting, producing visual features that are phonetically related. Next, we employ K-means to discretize the continuous audio and visual feature spaces. The discretization step allows DFVGM to capture high-level semantic structures that are more resilient to noise and generate visual hallucinations with high quality. To evaluate the effectiveness and robustness of our approach, we conduct extensive experiments on two publicly available datasets. The results demonstrate that our method achieves a remarkable 53% relative reduction (30.5%->12.9%) in Word Error Rate (WER) on average compared to the current state-of-the-art Audio-Only (AO) baselines while maintaining comparable results (< 5% difference) under the Audio-Visual (AV) setting even without video as input.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09ad17d4-275d-496a-89be-f715a59cd973Cited by top-tier papers5
- When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance CollapseYihuan Huang, Jun Xue, Liu Jiajun, Daixian Li et al.CVPR 2026 · 2 citations
- LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech RecognitionFeng Xue, Baochao Zhu, Wei Jia, Shujie Li et al.AAAI 2026
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang et al.ICLR 2025
- Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and ReactionHaoqiu Yan, Yongxin Zhu, Kai Zheng, Bing Liu et al.ACL 2024
- MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech RecognitionSungnyun Kim, Kangwook Jang, Sangmin Bae, Sungwoo Cho et al.ICML 2025
Builds on6
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech RecognitionXichen Pan, Peiyu Chen, Yichen Gong, Helong Zhou et al.ACL 2022 · 43 citations
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
- Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringJoanna Hong, Minsu Kim, Jeongsoo Choi, Yong Man RoCVPR 2023
Related papers
- AudioVSR: Enhancing Video Speech Recognition with Audio DataXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu et al.EMNLP 2024 · 3 citations
- Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech RecognitionYuchen Hu, Ruizhe Li, Chen Chen, Chengwei Qin et al.ACL 2023 · 7 citations
- Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and ManipulationChih-Chun Yang, Wan-Cyuan Fan, Cheng-Fu Yang, Yu-Chiang Frank WangAAAI 2022 · 16 citations
- Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation MappingYasser Abdelaziz Dahou Djilali, Sanath Narayan, Haithem Boussaid, Ebtesam Almazrouei et al.ICCV 2023 · 17 citations
- Restoring Speaking Lips from Occlusion for Audio-Visual Speech RecognitionJiadong Wang, Zexu Pan, Malu Zhang, Robby T. Tan et al.AAAI 2024 · 17 citations
