RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
Samuel Pegg, Kai Li, Xiaolin Hu
Abstract
Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA) models operate in the time domain. However, their overly simplistic approach to modeling acoustic features often necessitates larger and more computationally intensive models in order to achieve SOTA performance. In this paper, we present a novel time-frequency domain audio-visual speech separation method: Recurrent Time-Frequency Separation Network (RTFS-Net), which applies its algorithms on the complex time-frequency bins yielded by the Short-Time Fourier Transform. We model and capture the time and frequency dimensions of the audio independently using a multi-layered RNN along each dimension. Furthermore, we introduce a unique attention-based fusion technique for the efficient integration of audio and visual information, and a new mask separation approach that takes advantage of the intrinsic spectral nature of the acoustic features for a clearer separation. RTFS-Net outperforms the prior SOTA method in both inference speed and separation quality while reducing the number of parameters by 90% and MACs by 83%. This is the first time-frequency domain audio-visual speech separation method to outperform all contemporary time-domain counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5d4576a-6731-4cbb-9bdd-9113575a254dCited by top-tier papers5
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local AttentionKai Li, Gao Kejun, Xiaolin HuICLR 2026 · 5 citations
- RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual CuesTianrui Pan, Jie Liu, Bohan Wang, Jie Tang et al.ACM MM 2024 · 3 citations
- CogCM: Cognition-Inspired Contextual Modeling for Audio-Visual Speech EnhancementFeixiang Wang, Shuang Yang, Shiguang Shan, Xilin ChenICCV 2025 · 1 citation
- Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech SeparationXinmeng Xu, Haoran Xie, Xiaohui Tao, Lin Li et al.ICML 2026
- AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow MatchingXize Cheng, Chenyuhao Wen, Slytherin Wang, Yongqi Wang et al.ICLR 2026
Builds on5
- Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkXiaolin Hu, Kai Li, Weiyi Zhang, Yi Luo et al.NeurIPS 2021 · 74 citations
- The Right to Talk: An Audio-Visual Transformer ApproachThanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham et al.ICCV 2021 · 39 citations
- An efficient encoder-decoder architecture with top-down attention for speech separationKai Li, Runxuan Yang, Xiaolin HuICLR 2023 · 16 citations
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech SeparationJiyoung Lee, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang et al.CVPR 2021
Related papers
- IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationKai Li, Runxuan Yang, Fuchun Sun, Xiaolin HuICML 2024 · 28 citations
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan et al.AAAI 2021 · 112 citations
- Filter-Recovery Network for Multi-Speaker Audio-Visual Speech SeparationHaoyue Cheng, Zhaoyang Liu, Wayne Wu, Limin WangICLR 2023
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
- DTF-AT: Decoupled Time-Frequency Audio Transformer for Event ClassificationTony Alex, Sara Ahmed, Armin Mustafa, Muhammad Awais et al.AAAI 2024 · 8 citations
