Interactive Speech and Noise Modeling for Speech Enhancement
Chengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan, Yan Lu
摘要
Speech enhancement is challenging because of the diversity of background noise types. Most of the existing methods are focused on modelling the speech rather than the noise. In this paper, we propose a novel idea to model speech and noise simultaneously in a two-branch convolutional neural network, namely SN-Net. In SN-Net, the two branches predict speech and noise, respectively. Instead of information fusion only at the final output layer, interaction modules are introduced at several intermediate feature domains between the two branches to benefit each other. Such an interaction can leverage features learned from one branch to counteract the undesired part and restore the missing component of the other and thus enhance their discrimination capabilities. We also design a feature extraction module, namely residual-convolution-and-attention (RA), to capture the correlations along temporal and frequency dimensions for both the speech and the noises. Evaluations on public datasets show that the interaction module plays a key role in simultaneous modeling and the SN-Net outperforms the state-of-the-art by a large margin on various evaluation metrics. The proposed SN-Net also shows superior performance for speaker separation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Reference-Based Speech Enhancement via Feature Alignment and Fusion NetworkHuanjing Yue, Wenxin Duo, Xiulian Peng, Jingyu YangAAAI 2022 · 被引用 19 次
- Rethinking Flow and Diffusion Bridge Models for Speech EnhancementDahan Wang, Jun Gao, Tong Lei, Yuxiang Hu 等AAAI 2026 · 被引用 1 次
- Trainable EEG Interpolation and Structure-Sharing Dual-Path Encoders for Brain-Assisted Target Speaker ExtractionZhao Lv, Haoran Zhou, Ying Chen, Youdian Gao 等AAAI 2026
- D4AM: A General Denoising Framework for Downstream Acoustic ModelsChi-Chang Lee, Yu Tsao, Hsin-Min Wang, Chu-Song ChenICLR 2023
它引用的顶会 Paper2
相关 Paper
- FIRING-Net: A filtered feature recycling network for speech enhancementXinmeng Xu, Yiqun Zhang, Jizhen Li, Yuhong Yang 等ICLR 2025
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationSamuel Pegg, Kai Li, Xiaolin HuICLR 2024 · 被引用 13 次
- Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech EnhancementXinmeng Xu, Weiping Tu, Yuhong YangAAAI 2023 · 被引用 8 次
- IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationKai Li, Runxuan Yang, Fuchun Sun, Xiaolin HuICML 2024 · 被引用 28 次
- BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech EnhancementCunhang Fan, Enrui Liu, Andong Li, Jianhua Tao 等AAAI 2025
