Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech Enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang
Abstract
Attention mechanisms, such as local and non-local attention, play a fundamental role in recent deep learning based speech enhancement (SE) systems. However, natural speech contains many fast-changing and relatively brief acoustic events, therefore, capturing the most informative speech features by indiscriminately using local and non-local attention is challenged. We observe that the noise type and speech feature vary within a sequence of speech and the local and nonlocal operations can respectively extract different features from corrupted speech. To leverage this, we propose Selector-Enhancer, a dual-attention based convolution neural network (CNN) with a feature-filter that can dynamically select regions from low-resolution speech features and feed them to local or non-local attention operations. In particular, the proposed feature-filter is trained by using reinforcement learning (RL) with a developed difficulty-regulated reward that is related to network performance, model complexity, and "the difficulty of the SE task". The results show that our method achieves comparable or superior performance to existing approaches. In particular, Selector-Enhancer is potentially effective for real-world denoising, where the number and types of noise are varies on a single noisy mixture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 188963f4-8fac-409b-90ef-cd88220c587fCited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- FIRING-Net: A filtered feature recycling network for speech enhancementXinmeng Xu, Yiqun Zhang, Jizhen Li, Yuhong Yang et al.ICLR 2025
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan et al.AAAI 2021 · 112 citations
- SeDepTTS: Enhancing the Naturalness via Semantic Dependency and Local Convolution for Text-to-Speech SynthesisChenglong Jiang, Ying Gao, Wing W. Y. Ng, Jiyong Zhou et al.AAAI 2023 · 4 citations
- Deep Residual-Dense Lattice Network for Speech EnhancementMohammad Nikzad, Aaron Nicolson, Yongsheng Gao, Jun Zhou et al.AAAI 2020 · 42 citations
- Dual-view Attention Networks for Single Image Super-ResolutionJingcai Guo, Shiheng Ma, Jie Zhang, Qihua Zhou et al.ACM MM 2020 · 15 citations
