FIRING-Net: A filtered feature recycling network for speech enhancement
Xinmeng Xu, Yiqun Zhang, Jizhen Li, Yuhong Yang, Yong Luo, Weiping Tu
摘要
Current deep neural networks for speech enhancement (SE) aim to minimize the distance between the output signal and the clean target by filtering out noise features from input features. However, when noise and speech components are highly similar, SE models struggle to learn effective discrimination patterns. To address this challenge, we propose a Filter-Recycle-Interguide framework termed FIlter-Recycle-INterGuide NETwork (FIRING-Net) for SE, which filters the input features to extract target features and recycles the filtered-out features as non-target features. These two feature sets then guide each other to refine the features, leading to the aggregation of speech information within the target features and noise information within the non-target features. The proposed FIRING-Net mainly consists of a Local Module (LM) and a Global Module (GM). The LM uses outputs of the speech extraction network as target features and the residual between input and output as non-target features. The GM leverages the energy distribution of the self-attention map to extract target and non-target features guided by the highest and lowest energy regions. Both LM and GM include interaction modules to leverage the two feature sets in an inter-guided manner for collecting speech from non-target features and filtering out noise from target features. Experiments confirm the effectiveness of the Filter-Recycle-Interguide framework. Additionally, FIRING-Net achieves a good balance between SE performance and computational efficiency, outperforming other comparable models across various signalto-noise ratio levels and noise environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and UnderstandingYifan Peng, Siddharth Dalmia, Ian R. Lane, Shinji WatanabeICML 2022 · 被引用 203 次
- Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal TransportHsin-Yi Lin, Huan-Hsin Tseng, Xugang Lu, Yu TsaoNeurIPS 2021 · 被引用 40 次
- Learning A Sparse Transformer Network for Effective Image DerainingXiang Chen, Hao Li, Mingqiang Li, Jinshan PanCVPR 2023
相关 Paper
- Reference-Based Speech Enhancement via Feature Alignment and Fusion NetworkHuanjing Yue, Wenxin Duo, Xiulian Peng, Jingyu YangAAAI 2022 · 被引用 19 次
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan 等AAAI 2021 · 被引用 112 次
- Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech EnhancementXinmeng Xu, Weiping Tu, Yuhong YangAAAI 2023 · 被引用 8 次
- Deep Residual-Dense Lattice Network for Speech EnhancementMohammad Nikzad, Aaron Nicolson, Yongsheng Gao, Jun Zhou 等AAAI 2020 · 被引用 42 次
- Feature Decoupling-Recycling Network for Fast Interactive SegmentationHuimin Zeng, Weinong Wang, Xin Tao, Zhiwei Xiong 等ACM MM 2023 · 被引用 4 次
