Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
Sungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang, Se-Young Yun
摘要
Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely addressed audio disruptions, few studies have dealt with visual corruptions, e.g., lip occlusions or blurred videos, which are also detrimental. To address this real-world challenge, we propose CAV2vec, a novel self-supervised speech representation learning framework particularly designed to handle audio-visual joint corruption. CAV2vec employs a self-distillation approach with a corrupted prediction task, where the student model learns to predict clean targets, generated by the teacher model, with corrupted input frames. Specifically, we suggest a unimodal multi-task learning, which distills cross-modal knowledge and aligns the corrupted modalities, by predicting clean audio targets with corrupted videos, and clean video targets with corrupted audios. This strategy mitigates the dispersion in the representation space caused by corrupted modalities, leading to more reliable and robust audiovisual fusion. Our experiments on robust AVSR benchmarks demonstrate that the corrupted representation learning method significantly enhances recognition accuracy across generalized environments involving various types of corruption. Our code is available at https://github.com/sungnyun/cav2vec .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source LocalizationMin-Sang Baek, Gyeong-Su Kim, Donghyun Kim, Joon-Hyuk ChangICLR 2026
- MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech RecognitionSungnyun Kim, Kangwook Jang, Sangmin Bae, Sungwoo Cho 等ICML 2025
它引用的顶会 Paper23
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- Self-supervised learning with random-projection quantizer for speech recognitionChung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu 等ICML 2022 · 被引用 245 次
相关 Paper
- AV-RISE: Hierarchical Cross-Modal Denoising for Learning Robust Audio-Visual Speech RepresentationZhishuo Zhao, Yi Lin, Dongyue Guo, Junyu FanACM MM 2025 · 被引用 1 次
- Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringJoanna Hong, Minsu Kim, Jeongsoo Choi, Yong Man RoCVPR 2023
- Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech RepresentationQiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu 等AAAI 2024 · 被引用 17 次
- Enhancing Audio-Visual Association with Self-Supervised Curriculum LearningJingran Zhang, Xing Xu, Fumin Shen, Huimin Lu 等AAAI 2021 · 被引用 22 次
- AV-TranSpeech: Audio-Visual Robust Speech-to-Speech TranslationRongjie Huang, Huadai Liu, Xize Cheng, Yi Ren 等ACL 2023 · 被引用 9 次
