Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
Qiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu, Lirong Dai
Abstract
Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the scarcity of labeled multichannel data and complex ambient noises. The efficacy of self-supervised learning for far-field multichannel and multi-modal speech processing has not been well explored. Considering that visual information helps to improve speech recognition performance in noisy scenes, in this work we propose the multichannel multi-modal speech self-supervised learning framework AV-wav2vec2, which utilizes video and multichannel audio data as inputs. First, we propose a multi-path structure to process multi-channel audio streams and a visual stream in parallel, with intra-, and inter-channel contrastive as training targets to fully exploit the rich information in multi-channel speech data. Second, based on contrastive learning, we use additional single-channel audio data, which is trained jointly to improve the performance of multichannel multi-modal representation. Finally, we use a Chinese multichannel multi-modal dataset in real scenarios to validate the effectiveness of the proposed method on audio-visual speech recognition (AVSR), automatic speech recognition (ASR), visual speech recognition (VSR) and audio-visual speaker diarization (AVSD) tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7efebebd-e187-49a9-9605-11d55c89acdcCited by top-tier papers2
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang et al.ICLR 2025
- Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source LocalizationMin-Sang Baek, Gyeong-Su Kim, Donghyun Kim, Joon-Hyuk ChangICLR 2026
Builds on3
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech RecognitionYuchen Hu, Chen Chen, Ruizhe Li, Heqing Zou et al.ACL 2023 · 12 citations
- Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech RecognitionYuchen Hu, Ruizhe Li, Chen Chen, Chengwei Qin et al.ACL 2023 · 7 citations
Related papers
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath et al.ICLR 2023 · 17 citations
- Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech RecognitionXichen Pan, Peiyu Chen, Yichen Gong, Helong Zhou et al.ACL 2022 · 43 citations
- MAViL: Masked Audio-Video LearnersPo-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali et al.NeurIPS 2023 · 95 citations
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er et al.AAAI 2023 · 13 citations
- Enhancing Audio-Visual Association with Self-Supervised Curriculum LearningJingran Zhang, Xing Xu, Fumin Shen, Huimin Lu et al.AAAI 2021 · 22 citations
