Not Only Vision: Evolve Visual Speech Recognition via Peripheral Information
Zhaoxin Yuan, Shuang Yang, Shiguang Shan, Xilin Chen
Abstract
Is visual information alone sufficient for visual speech recognition (VSR) in challenging real-world scenarios? Humans do not rely solely on visual information for lipreading but also incorporate additional cues, such as speech-related context and prior knowledge about the task. However, existing methods have largely overlooked such external information in automatic VSR systems. To systematically explore the role of such information for VSR, we introduce the concept of Peripheral Information. We categorize it into three types based on the relevance to the spoken content: (1) Contextual Guidance (e.g., topic or description of speech), (2) Task Expertise (e.g., human prior experience in lip-reading), and (3) Linguistic Perturbation (irrelevant signals processed alongside meaningful information). Considering the disparity that peripheral information provides additional clues with varying significance while visual input serves as the most direct source for VSR, we propose a framework that introduces a hierarchical processing strategy to handle different modalities. With visualspecific adaptation and a dynamic routing mechanism for multi-modal information, our approach reduces the impact of modality conflicts effectively and enables selective utilization of peripheral information with varying relevance. Leveraging readily available peripheral information, our model achieves a WER of 22.03% on LRS3. Further experiments on AVSpeech demonstrate its generalization in realworld scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c7c0880-3ee9-41f0-b133-ee8f42e117edBuilds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-TuningChunlin Tian, Zhan Shi, Zhijiang Guo, Li Li et al.NeurIPS 2024 · 172 citations
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 104 citations
Related papers
- Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringJoanna Hong, Minsu Kim, Jeongsoo Choi, Yong Man RoCVPR 2023
- Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation MappingYasser Abdelaziz Dahou Djilali, Sanath Narayan, Haithem Boussaid, Ebtesam Almazrouei et al.ICCV 2023 · 17 citations
- Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement LearningChen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou et al.AAAI 2023 · 35 citations
- VALLR: Visual ASR Language Model for Lip ReadingMarshall Thomas, Edward Fish, Richard BowdenICCV 2025 · 6 citations
- Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip ReadingSucheng Ren, Yong Du, Jianming Lv, Guoqiang Han et al.CVPR 2021
