IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
Kai Li, Runxuan Yang, Fuchun Sun, Xiaolin Hu
摘要
Recent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual features without employing selective attention mechanisms, which is in sharp contrast with the brain. To address this issue, We propose a novel model called Intra- and Inter-Attention Network (IIANet), which leverages the attention mechanism for efficient audio-visual feature fusion. IIANet consists of two types of attention blocks: intra-attention (IntraA) and inter-attention (InterA) blocks, where the InterA blocks are distributed at the top, middle and bottom of IIANet. Heavily inspired by the way how human brain selectively focuses on relevant content at various temporal scales, these blocks maintain the ability to learn modality-specific features and enable the extraction of different semantics from audio-visual features. Comprehensive experiments on three standard audio-visual separation benchmarks (LRS2, LRS3, and VoxCeleb2) demonstrate the effectiveness of IIANet, outperforming previous state-of-the-art methods while maintaining comparable inference time. In particular, the fast version of IIANet (IIANet-fast) has only 7% of CTCNet's MACs and is 40% faster than CTCNet on CPUs while achieving better separation quality, showing the great potential of attention mechanism for efficient and effective multimodal fusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SAM Audio: Segment Anything in AudioBowen Shi, Andros Tjandra, John Hoffman, Helin Wang 等ICML 2026 · 被引用 35 次
- Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural ConversationTaekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon 等CVPR 2026 · 被引用 18 次
- Dynamic Dictionary Learning for Remote Sensing Image SegmentationXuechao Zou, Yue Li, Shun Zhang, Kai Li 等ICCV 2025 · 被引用 15 次
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local AttentionKai Li, Gao Kejun, Xiaolin HuICLR 2026 · 被引用 5 次
- CogCM: Cognition-Inspired Contextual Modeling for Audio-Visual Speech EnhancementFeixiang Wang, Shuang Yang, Shiguang Shan, Xilin ChenICCV 2025 · 被引用 1 次
它引用的顶会 Paper5
- Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkXiaolin Hu, Kai Li, Weiyi Zhang, Yi Luo 等NeurIPS 2021 · 被引用 74 次
- Reading to Listen at the Cocktail Party: Multi-Modal Speech SeparationAkam Rahimi, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 21 次
- An efficient encoder-decoder architecture with top-down attention for speech separationKai Li, Runxuan Yang, Xiaolin HuICLR 2023 · 被引用 16 次
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech SeparationJiyoung Lee, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang 等CVPR 2021
相关 Paper
- Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech SeparationXinmeng Xu, Haoran Xie, Xiaohui Tao, Lin Li 等ICML 2026
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationSamuel Pegg, Kai Li, Xiaolin HuICLR 2024 · 被引用 13 次
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 等ACM MM 2021 · 被引用 154 次
- Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationHanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang 等AAAI 2020 · 被引用 110 次
- Audio-Visual Glance Network for Efficient Video RecognitionMuhammad Adi Nugroho, Sangmin Woo, Sumin Lee, Changick KimICCV 2023 · 被引用 8 次
