MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition
Yuchen Hu, Chen Chen, Ruizhe Li, Heqing Zou, Eng Siong Chng
Abstract
Audio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech. Mainstream approaches addressing this task have developed sophisticated architectures and techniques for multi-modality fusion and representation learning. However, the natural heterogeneity of different modalities causes distribution gap between their representations, making it challenging to fuse them. In this paper, we aim to learn the shared representations across modalities to bridge their gap. Different from existing similar methods on other multimodal tasks like sentiment analysis, we focus on the temporal contextual dependencies considering the sequence-to-sequence task setting of AVSR. In particular, we propose an adversarial network to refine framelevel modality-invariant representations (MIR-GAN), which captures the commonality across modalities to ease the subsequent multimodal fusion process. Extensive experiments on public benchmarks LRS3 and LRS2 show that our approach outperforms the state-of-the-arts 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech RepresentationQiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu et al.AAAI 2024 · 17 citations
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang et al.ICLR 2025
Builds on10
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Cross-Modality Person Re-Identification via Modality Confusion and Center AggregationXin Hao, Sanyuan Zhao, Mang Ye, Jianbing ShenICCV 2021 · 191 citations
Related papers
- Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and ManipulationChih-Chun Yang, Wan-Cyuan Fan, Cheng-Fu Yang, Yu-Chiang Frank WangAAAI 2022 · 16 citations
- Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement LearningChen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou et al.AAAI 2023 · 35 citations
- Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionSijie Mai, Haifeng Hu, Songlong XingAAAI 2020 · 233 citations
- Audio-Visual Semantic Graph Network for Audio-Visual Event LocalizationLiang Liu, Shuaiyong Li, Yongqiang ZhuCVPR 2025
- AV-RISE: Hierarchical Cross-Modal Denoising for Learning Robust Audio-Visual Speech RepresentationZhishuo Zhao, Yi Lin, Dongyue Guo, Junyu FanACM MM 2025 · 1 citation
