MSFNet: Multi-Scale Fusion Network for Brain-Controlled Speaker Extraction
Cunhang Fan, Jingjing Zhang, Hongyu Zhang, Wang Xiang, Jianhua Tao, Xinhui Li, Jiangyan Yi, Dianbo Sui, Zhao Lv
Abstract
Speaker extraction aims to selectively extract the target speaker from the multi-talker environment under the guidance of auxiliary reference. Recent studies have shown that the attended speaker's information can be decoded by the auditory attention decoding from the listener's brain activity. However, how to more effectively utilize the common information about the target speaker contained in both electroencephalography (EEG) and speech is still an unresolved problem. In this paper, we propose a multi-scale fusion network (MSFNet) for brain-controlled speaker extraction, which utilizes the EEG recorded from the listener to extract the target speech. In order to make full use of the speech information, the mixed speech is encoded with multiple time scales so that the multi-scale embeddings are acquired. In addition, to effectively extract the non-Euclidean data of EEG, the graph convolutional networks are used as the EEG encoder. Finally, these multi-scale embeddings are separately fused with the EEG features. To facilitate research related to auditory attention decoding and further validate the effectiveness of the proposed method, we also construct the AVED dataset, a new EEG-Audio dataset. Experimental results on both the public Cocktail Party dataset and the newly proposed AVED dataset in this paper show that our MSFNet model significantly outperforms the state-of-the-art method in certain objective evaluation metrics.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d2c5e7ca-a3a9-47e0-ba11-e07e6d8bbc62Cited by top-tier papers3
- DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram ReconstructionCunhang Fan, Sheng Zhang, Jingjing Zhang, Enrui Liu et al.ACM MM 2025 · 4 citations
- Trainable EEG Interpolation and Structure-Sharing Dual-Path Encoders for Brain-Assisted Target Speaker ExtractionZhao Lv, Haoran Zhou, Ying Chen, Youdian Gao et al.AAAI 2026
- PCRNet: Phase-aware Complex Refinement Network for EEG-based Auditory Attention DecodingXiran Chen, Xiaoke Yang, Jian Zhou, Zhao Lv et al.ICML 2026
Related papers
- DHGCN: Dual HyperGraph Convolutional Network for EEG-Based Auditory Attention DetectionJian Zhou, Yingjie Xie, Cunhang Fan, Huabin Wang et al.ACM MM 2025 · 5 citations
- Auditory Attention Decoding with Task-Related Multi-View Contrastive LearningXiaoyu Chen, Changde Du, Qiongyi Zhou, Huiguang HeACM MM 2023 · 5 citations
- DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention DetectionSheng Yan, Cunhang Fan, Hongyu Zhang, Xiaoke Yang et al.NeurIPS 2024 · 57 citations
- MAAS: Multi-modal Assignation for Active Speaker DetectionJuan León Alcázar, Fabian Caba Heilbron, Ali K. Thabet, Bernard GhanemICCV 2021 · 66 citations
- A Speaker-Aware Co-Attention Framework for Medical Dialogue Information ExtractionYuan Xia, Zhenhui Shi, Jingbo Zhou, Jiayu Xu et al.EMNLP 2022 · 6 citations
