Restoring Speaking Lips from Occlusion for Audio-Visual Speech Recognition
Jiadong Wang, Zexu Pan, Malu Zhang, Robby T. Tan, Haizhou Li
摘要
Prior studies on audio-visual speech recognition typically assume the visibility of speaking lips, ignoring the fact that visual occlusion occurs in real-world videos, thus adversely affecting recognition performance. To address this issue, we propose a framework that restores occluded lips in a video by utilizing both the video itself and the corresponding noisy audio. Specifically, the framework aims to achieve these three tasks: detecting occluded frames, masking occluded areas, and reconstruction of masked regions. We tackle the first two issues by utilizing the Class Activation Mapping (CAM) obtained from occluded frame detection to facilitate the masking of occluded areas. Additionally, we introduce a novel synthesis-matching strategy for the reconstruction to ensure the compatibility of audio features with different levels of occlusion. Our framework is evaluated in terms of Word Error Rate (WER) on the original videos, the videos corrupted by concealed lips, and the videos restored using the framework with several existing state-of-the-art audiovisual speech recognition methods. Experimental results substantiate that our framework significantly mitigates performance degradation resulting from lip occlusion. Under -5dB noise conditions, AV-Hubert's WER increases from 10.62% to 13.87% due to lip occlusion, but rebounds to 11.87% in conjunction with the proposed framework. Furthermore, the framework also demonstrates its capacity to produce natural synthesized images in qualitative assessments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual CuesTianrui Pan, Jie Liu, Bohan Wang, Jie Tang 等ACM MM 2024 · 被引用 3 次
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang 等ICLR 2025
- MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech RecognitionSungnyun Kim, Kangwook Jang, Sangmin Bae, Sungwoo Cho 等ICML 2025
它引用的顶会 Paper8
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- DC-ShadowNet: Single-Image Hard and Soft Shadow Removal Using Unsupervised Domain-Classifier Guided NetworkYeying Jin, Aashish Sharma, Robby T. TanICCV 2021 · 被引用 163 次
- Expressive Talking Head Generation with Granular Audio-Visual ControlBorong Liang, Yan Pan, Zhizhi Guo, Hang Zhou 等CVPR 2022 · 被引用 114 次
- SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip MemorySe Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi 等AAAI 2022 · 被引用 110 次
相关 Paper
- Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringJoanna Hong, Minsu Kim, Jeongsoo Choi, Yong Man RoCVPR 2023
- Visual Hallucination Elevates Speech RecognitionFang Zhang, Yongxin Zhu, Xiangxiang Wang, Huang Chen 等AAAI 2024 · 被引用 5 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoXiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu 等ICCV 2023 · 被引用 18 次
- Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation MappingYasser Abdelaziz Dahou Djilali, Sanath Narayan, Haithem Boussaid, Ebtesam Almazrouei 等ICCV 2023 · 被引用 17 次
