IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
Kai Li, Runxuan Yang, Fuchun Sun, Xiaolin Hu
Abstract
Recent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual features without employing selective attention mechanisms, which is in sharp contrast with the brain. To address this issue, We propose a novel model called Intra- and Inter-Attention Network (IIANet), which leverages the attention mechanism for efficient audio-visual feature fusion. IIANet consists of two types of attention blocks: intra-attention (IntraA) and inter-attention (InterA) blocks, where the InterA blocks are distributed at the top, middle and bottom of IIANet. Heavily inspired by the way how human brain selectively focuses on relevant content at various temporal scales, these blocks maintain the ability to learn modality-specific features and enable the extraction of different semantics from audio-visual features. Comprehensive experiments on three standard audio-visual separation benchmarks (LRS2, LRS3, and VoxCeleb2) demonstrate the effectiveness of IIANet, outperforming previous state-of-the-art methods while maintaining comparable inference time. In particular, the fast version of IIANet (IIANet-fast) has only 7% of CTCNet's MACs and is 40% faster than CTCNet on CPUs while achieving better separation quality, showing the great potential of attention mechanism for efficient and effective multimodal fusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8da2828b-3b9d-4897-b6d5-c22a7b87f0c5Cited by top-tier papers6
- SAM Audio: Segment Anything in AudioBowen Shi, Andros Tjandra, John Hoffman, Helin Wang et al.ICML 2026 · 35 citations
- Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural ConversationTaekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon et al.CVPR 2026 · 18 citations
- Dynamic Dictionary Learning for Remote Sensing Image SegmentationXuechao Zou, Yue Li, Shun Zhang, Kai Li et al.ICCV 2025 · 15 citations
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local AttentionKai Li, Gao Kejun, Xiaolin HuICLR 2026 · 5 citations
- CogCM: Cognition-Inspired Contextual Modeling for Audio-Visual Speech EnhancementFeixiang Wang, Shuang Yang, Shiguang Shan, Xilin ChenICCV 2025 · 1 citation
Builds on5
- Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkXiaolin Hu, Kai Li, Weiyi Zhang, Yi Luo et al.NeurIPS 2021 · 74 citations
- Reading to Listen at the Cocktail Party: Multi-Modal Speech SeparationAkam Rahimi, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 21 citations
- An efficient encoder-decoder architecture with top-down attention for speech separationKai Li, Runxuan Yang, Xiaolin HuICLR 2023 · 16 citations
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech SeparationJiyoung Lee, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang et al.CVPR 2021
Related papers
- Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech SeparationXinmeng Xu, Haoran Xie, Xiaohui Tao, Lin Li et al.ICML 2026
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationSamuel Pegg, Kai Li, Xiaolin HuICLR 2024 · 13 citations
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian et al.ACM MM 2021 · 154 citations
- Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationHanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang et al.AAAI 2020 · 110 citations
- Audio-Visual Glance Network for Efficient Video RecognitionMuhammad Adi Nugroho, Sangmin Woo, Sumin Lee, Changick KimICCV 2023 · 8 citations
