FIRING-Net: A filtered feature recycling network for speech enhancement
Xinmeng Xu, Yiqun Zhang, Jizhen Li, Yuhong Yang, Yong Luo, Weiping Tu
Abstract
Current deep neural networks for speech enhancement (SE) aim to minimize the distance between the output signal and the clean target by filtering out noise features from input features. However, when noise and speech components are highly similar, SE models struggle to learn effective discrimination patterns. To address this challenge, we propose a Filter-Recycle-Interguide framework termed FIlter-Recycle-INterGuide NETwork (FIRING-Net) for SE, which filters the input features to extract target features and recycles the filtered-out features as non-target features. These two feature sets then guide each other to refine the features, leading to the aggregation of speech information within the target features and noise information within the non-target features. The proposed FIRING-Net mainly consists of a Local Module (LM) and a Global Module (GM). The LM uses outputs of the speech extraction network as target features and the residual between input and output as non-target features. The GM leverages the energy distribution of the self-attention map to extract target and non-target features guided by the highest and lowest energy regions. Both LM and GM include interaction modules to leverage the two feature sets in an inter-guided manner for collecting speech from non-target features and filtering out noise from target features. Experiments confirm the effectiveness of the Filter-Recycle-Interguide framework. Additionally, FIRING-Net achieves a good balance between SE performance and computational efficiency, outperforming other comparable models across various signalto-noise ratio levels and noise environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and UnderstandingYifan Peng, Siddharth Dalmia, Ian R. Lane, Shinji WatanabeICML 2022 · 203 citations
- Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal TransportHsin-Yi Lin, Huan-Hsin Tseng, Xugang Lu, Yu TsaoNeurIPS 2021 · 40 citations
- Learning A Sparse Transformer Network for Effective Image DerainingXiang Chen, Hao Li, Mingqiang Li, Jinshan PanCVPR 2023
Related papers
- Reference-Based Speech Enhancement via Feature Alignment and Fusion NetworkHuanjing Yue, Wenxin Duo, Xiulian Peng, Jingyu YangAAAI 2022 · 19 citations
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan et al.AAAI 2021 · 112 citations
- Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech EnhancementXinmeng Xu, Weiping Tu, Yuhong YangAAAI 2023 · 8 citations
- Deep Residual-Dense Lattice Network for Speech EnhancementMohammad Nikzad, Aaron Nicolson, Yongsheng Gao, Jun Zhou et al.AAAI 2020 · 42 citations
- Feature Decoupling-Recycling Network for Fast Interactive SegmentationHuimin Zeng, Weinong Wang, Xin Tao, Zhiwei Xiong et al.ACM MM 2023 · 4 citations
