Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement Learning
Chen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou, Beier Zhu, Eng Siong Chng
摘要
Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant representations. However, such representations are prone to over-reliance on audio modality as it is much easier to recognize than video modality in clean conditions. As a result, the AVSR model underestimates the importance of visual stream in face of noise corruption. To this end, we leverage visual modality-specific representations to provide stable complementary information for the AVSR task. Specifically, we propose a reinforcement learning (RL) based framework called MSRL, where the agent dynamically harmonizes modality-invariant and modality-specific representations in the auto-regressive decoding process. We customize a reward function directly related to task-specific metrics (i.e., word error rate), which encourages the MSRL to effectively explore the optimal integration strategy. Experimental results on the LRS3 dataset show that the proposed method achieves state-of-the-art in both clean and various noisy conditions. Furthermore, we demonstrate the better generality of MSRL system than other baselines when test set contains unseen noises.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech RecognitionChen Chen, Ruizhe Li, Yuchen Hu, Sabato Marco Siniscalchi 等ICLR 2024 · 被引用 37 次
- Neural Oscillators for Generalization of Physics-Informed Machine LearningTaniya Kapoor, Abhishek Chandra, Daniel M. Tartakovsky, Hongrui Wang 等AAAI 2024 · 被引用 17 次
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech RecognitionUmberto Cappellazzo, Minsu Kim, Pingchuan Ma, Honglie Chen 等NeurIPS 2025 · 被引用 5 次
- A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech RecognitionYusheng Dai, Hang Chen, Jun Du, Ruoyu Wang 等CVPR 2024
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang 等ICLR 2025
它引用的顶会 Paper5
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- M3ER: Multiplicative Multimodal Emotion Recognition using Facial, Textual, and Speech CuesTrisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera 等AAAI 2020 · 被引用 282 次
- Reinforcement Learning Based Dynamic Model Combination for Time Series ForecastingYuwei Fu, Di Wu, Benoit BouletAAAI 2022 · 被引用 66 次
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
相关 Paper
- Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringJoanna Hong, Minsu Kim, Jeongsoo Choi, Yong Man RoCVPR 2023
- AV-RISE: Hierarchical Cross-Modal Denoising for Learning Robust Audio-Visual Speech RepresentationZhishuo Zhao, Yi Lin, Dongyue Guo, Junyu FanACM MM 2025 · 被引用 1 次
- Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech RecognitionYuchen Hu, Ruizhe Li, Chen Chen, Chengwei Qin 等ACL 2023 · 被引用 7 次
- Jointly Learning Visual and Auditory Speech Representations from Raw DataAlexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis 等ICLR 2023 · 被引用 13 次
- Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and ManipulationChih-Chun Yang, Wan-Cyuan Fan, Cheng-Fu Yang, Yu-Chiang Frank WangAAAI 2022 · 被引用 16 次
