BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement
Cunhang Fan, Enrui Liu, Andong Li, Jianhua Tao, Jian Zhou, Jiahao Li, Chengshi Zheng, Zhao Lv
摘要
Although the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the performance of SE, many modules are stacked onto SE, resulting in increased model complexity that limits the application of SE. To address these problems, we proposed a dual-path network based on compressed frequency using Mamba. First, we extract amplitude and phase information through parallel dual branches. This approach leverages structured complex spectra to implicitly capture phase information and solves the compensation effect by decoupling amplitude and phase, and the network incorporates an interaction module to suppress unnecessary parts and recover missing components from the other branch. Second, to reduce network complexity, the network introduces a band-split strategy to compress the frequency dimension. To further reduce complexity while maintaining good performance, we designed a Mamba-based module that models the time and frequency dimensions under linear complexity. Finally, compared to baselines, our model achieves an average 8.3 times reduction in computational complexity while maintaining superior performance. Furthermore, it achieves a 25 times reduction in complexity compared to transformer-based models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Robust Signal Enhancement via Fractional Detail Views and Knowledge Guided Multi-view FusionZikun Jin, Yuhua Qian, Xinyan Liang, Jiaqian Zhang 等ICML 2026
- Signal Enhancement via Multi-view Dynamic Representation and Alignment-aware FusionZikun Jin, Yuhua Qian, Xinyan Liang, Jiaqian Zhang 等AAAI 2026
它引用的顶会 Paper6
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra 等NeurIPS 2020 · 被引用 1,100 次
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 被引用 387 次
- Reference-Based Speech Enhancement via Feature Alignment and Fusion NetworkHuanjing Yue, Wenxin Duo, Xiulian Peng, Jingyu YangAAAI 2022 · 被引用 19 次
- Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech EnhancementXinmeng Xu, Weiping Tu, Yuhong YangAAAI 2023 · 被引用 8 次
相关 Paper
- VmambaSCI: Dynamic Deep Unfolding Network with Mamba for Compressive Spectral ImagingMingjin Zhang, Longyi Li, Wenxuan Shi, Jie Guo 等ACM MM 2024 · 被引用 13 次
- HiFi-Mamba: Dual-Stream ?-Laplacian Enhanced Mamba for High-Fidelity MRI ReconstructionHongli Chen, Pengcheng Fang, Yuxia Chen, Yingxuan Ren 等AAAI 2026 · 被引用 2 次
- Complex-Cycle-Consistent Diffusion Model for Monaural Speech EnhancementYi Li, Yang Sun, Plamen P. AngelovAAAI 2025 · 被引用 2 次
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan 等AAAI 2021 · 被引用 112 次
- TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable PlatformsYueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang 等UbiComp 2025 · 被引用 18 次
