Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
Kai Li, Gao Kejun, Xiaolin Hu
摘要
Audio-visual speech separation (AVSS) methods leverage visual cues to extract target speech and have demonstrated strong separation quality in noisy acoustic environments. However, these methods usually involve a large number of parameters and require high computational cost, which is unacceptable in many applications where speech separation serves as only a preprocessing step for further speech processing. To address this issue, we propose an efficient AVSS method, named Dolphin. For visual feature extraction, we develop DP‑LipCoder, a dual‑path lightweight video encoder that transforms lip‑motion into discrete audio‑aligned semantic tokens. For audio separation, we construct a lightweight encoder–decoder separator, in which each layer incorporates a global–local attention (GLA) block to efficiently capture multi-scale dependencies. Experiments on three benchmark datasets showed that Dolphin not only surpassed the current state-of-the-art (SOTA) model in separation quality but also achieved remarkable improvements in efficiency: over 50% fewer parameters, more than 2.4 reduction in MACs, and over 6 faster GPU inference speed. These results indicate that Dolphin offers a practical and deployable solution for high-performance AVSS in real-world scenarios. Our code and demo page are publicly available at https://cslikai.cn/Dolphin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech SeparationUi-Hyeop Shin, Sangyoun Lee, Taehan Kim, Hyung-Min ParkNeurIPS 2024 · 被引用 46 次
- IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationKai Li, Runxuan Yang, Fuchun Sun, Xiaolin HuICML 2024 · 被引用 28 次
- An efficient encoder-decoder architecture with top-down attention for speech separationKai Li, Runxuan Yang, Xiaolin HuICLR 2023 · 被引用 16 次
相关 Paper
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- LightAVSeg: Lightweight Audio-Visual SegmentationQing Zhong, Guodong Ding, Lingqiao Liu, Zaiwen Feng 等ICML 2026
- Audio-Visual Glance Network for Efficient Video RecognitionMuhammad Adi Nugroho, Sangmin Woo, Sumin Lee, Changick KimICCV 2023 · 被引用 8 次
- Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image SynergyKe Xue, Rongfei Fan, Kai Li, Shanping Yu 等ICML 2026
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationSamuel Pegg, Kai Li, Xiaolin HuICLR 2024 · 被引用 13 次
