RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
Samuel Pegg, Kai Li, Xiaolin Hu
摘要
Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA) models operate in the time domain. However, their overly simplistic approach to modeling acoustic features often necessitates larger and more computationally intensive models in order to achieve SOTA performance. In this paper, we present a novel time-frequency domain audio-visual speech separation method: Recurrent Time-Frequency Separation Network (RTFS-Net), which applies its algorithms on the complex time-frequency bins yielded by the Short-Time Fourier Transform. We model and capture the time and frequency dimensions of the audio independently using a multi-layered RNN along each dimension. Furthermore, we introduce a unique attention-based fusion technique for the efficient integration of audio and visual information, and a new mask separation approach that takes advantage of the intrinsic spectral nature of the acoustic features for a clearer separation. RTFS-Net outperforms the prior SOTA method in both inference speed and separation quality while reducing the number of parameters by 90% and MACs by 83%. This is the first time-frequency domain audio-visual speech separation method to outperform all contemporary time-domain counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local AttentionKai Li, Gao Kejun, Xiaolin HuICLR 2026 · 被引用 5 次
- RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual CuesTianrui Pan, Jie Liu, Bohan Wang, Jie Tang 等ACM MM 2024 · 被引用 3 次
- CogCM: Cognition-Inspired Contextual Modeling for Audio-Visual Speech EnhancementFeixiang Wang, Shuang Yang, Shiguang Shan, Xilin ChenICCV 2025 · 被引用 1 次
- Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech SeparationXinmeng Xu, Haoran Xie, Xiaohui Tao, Lin Li 等ICML 2026
- AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow MatchingXize Cheng, Chenyuhao Wen, Slytherin Wang, Yongqi Wang 等ICLR 2026
它引用的顶会 Paper5
- Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkXiaolin Hu, Kai Li, Weiyi Zhang, Yi Luo 等NeurIPS 2021 · 被引用 74 次
- The Right to Talk: An Audio-Visual Transformer ApproachThanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham 等ICCV 2021 · 被引用 39 次
- An efficient encoder-decoder architecture with top-down attention for speech separationKai Li, Runxuan Yang, Xiaolin HuICLR 2023 · 被引用 16 次
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech SeparationJiyoung Lee, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang 等CVPR 2021
相关 Paper
- IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationKai Li, Runxuan Yang, Fuchun Sun, Xiaolin HuICML 2024 · 被引用 28 次
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan 等AAAI 2021 · 被引用 112 次
- Filter-Recovery Network for Multi-Speaker Audio-Visual Speech SeparationHaoyue Cheng, Zhaoyang Liu, Wayne Wu, Limin WangICLR 2023
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
- DTF-AT: Decoupled Time-Frequency Audio Transformer for Event ClassificationTony Alex, Sara Ahmed, Armin Mustafa, Muhammad Awais 等AAAI 2024 · 被引用 8 次
