Reference-Based Speech Enhancement via Feature Alignment and Fusion Network
Huanjing Yue, Wenxin Duo, Xiulian Peng, Jingyu Yang
Abstract
Speech enhancement aims at recovering a clean speech from a noisy input, which can be classified into single speech enhancement and personalized speech enhancement. Personalized speech enhancement usually utilizes the speaker identity extracted from the noisy speech itself (or a clean reference speech) as a global embedding to guide the enhancement process. Different from them, we observe that the speeches of the same speaker are correlated in terms of frame-level short-time Fourier Transform (STFT) spectrogram. Therefore, we propose reference-based speech enhancement via a feature alignment and fusion network (FAF-Net). Given a noisy speech and a clean reference speech spoken by the same speaker, we first propose a feature-level alignment strategy to warp the clean reference with the noisy speech in frame level. Then, we fuse the reference feature with the noisy feature via a similarity-based fusion strategy. Finally, the fused features are skipped connected to the decoder, which generates the enhanced results. Experimental results demonstrate that the performance of the proposed FAF-Net is close to the state-of-the-art speech enhancement methods on both DNS and Voice Bank+DEMAND datasets. Our code is available at https://github.com/HieDean/FAF-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04720126-dfd1-412b-9e61-66af22db716fCited by top-tier papers2
- BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech EnhancementCunhang Fan, Enrui Liu, Andong Li, Jianhua Tao et al.AAAI 2025
- D4AM: A General Denoising Framework for Downstream Acoustic ModelsChi-Chang Lee, Yu Tsao, Hsin-Min Wang, Chu-Song ChenICLR 2023
Builds on3
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 387 citations
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan et al.AAAI 2021 · 112 citations
- Learning Texture Transformer Network for Image Super-ResolutionFuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu et al.CVPR 2020
Related papers
- FIRING-Net: A filtered feature recycling network for speech enhancementXinmeng Xu, Yiqun Zhang, Jizhen Li, Yuhong Yang et al.ICLR 2025
- Zero-Shot Face-Based Voice Conversion: Bottleneck-Free Speech Disentanglement in the Real-World ScenarioShao-En Weng, Hong-Han Shuai, Wen-Huang ChengAAAI 2023 · 4 citations
- Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-SynthesisKarren Yang, Dejan Markovic, Steven Krenn, Vasu Agrawal et al.CVPR 2022 · 32 citations
- UltraSpeech: Speech Enhancement by Interaction between Ultrasound and SpeechHan Ding, Yizhan Wang, Hao Li, Cui Zhao et al.UbiComp 2022 · 32 citations
- Robust Reference-Based Super-Resolution With Similarity-Aware Deformable ConvolutionGyumin Shim, Jinsun Park, In So KweonCVPR 2020
