Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman Mohamed
摘要
Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker's lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT (AV-HuBERT), a self-supervised representation learning framework for audio-visual speech, which masks multi-stream video input and predicts automatically discovered and iteratively refined multimodal hidden units. AV-HuBERT learns powerful audio-visual speech representation benefiting both lip-reading and automatic speech recognition. On the largest public lip-reading benchmark LRS3 (433 hours), AV-HuBERT achieves 32.5% WER with only 30 hours of labeled data, outperforming the former state-of-the-art approach (33.6%) trained with a thousand times more transcribed video data (31K hours) (Makino et al., 2019) . The lip-reading WER is further reduced to 26.9% when using all 433 hours of labeled data from LRS3 and combined with self-training. Using our audio-visual representation on the same benchmark for audio-only speech recognition leads to a 40% relative WER reduction over the state-of-the-art performance (1.3% vs 2.3%). Our code and models are available at https://github.com/ facebookresearch/av_hubert
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper77
- MAViL: Masked Audio-Video LearnersPo-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali 等NeurIPS 2023 · 被引用 95 次
- u-HuBERT: Unified Mixed-Modal Speech Pretraining And Zero-Shot Transfer to Unlabeled ModalityWei-Ning Hsu, Bowen ShiNeurIPS 2022 · 被引用 58 次
- TVLT: Textless Vision-Language TransformerZineng Tang, Jaemin Cho, Yixin Nie, Mohit BansalNeurIPS 2022 · 被引用 40 次
- It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech RecognitionChen Chen, Ruizhe Li, Yuchen Hu, Sabato Marco Siniscalchi 等ICLR 2024 · 被引用 37 次
- Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement LearningChen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou 等AAAI 2023 · 被引用 35 次
它引用的顶会 Paper7
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani 等NeurIPS 2020 · 被引用 483 次
- Parameter Efficient Multimodal Transformers for Video Representation LearningSangho Lee, Youngjae Yu, Gunhee Kim, Thomas M. Breuel 等ICLR 2021 · 被引用 90 次
- Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingXingxuan Zhang, Feng Cheng, Shilin WangICCV 2019 · 被引用 87 次
相关 Paper
- SynthVSR: Scaling Up Visual Speech RecognitionWith Synthetic SupervisionXubo Liu, Egor Lakomkin, Konstantinos Vougioukas, Pingchuan Ma 等CVPR 2023
- VALLR: Visual ASR Language Model for Lip ReadingMarshall Thomas, Edward Fish, Richard BowdenICCV 2025 · 被引用 6 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation MappingYasser Abdelaziz Dahou Djilali, Sanath Narayan, Haithem Boussaid, Ebtesam Almazrouei 等ICCV 2023 · 被引用 17 次
- Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech RecognitionXichen Pan, Peiyu Chen, Yichen Gong, Helong Zhou 等ACL 2022 · 被引用 43 次
