DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
Cunhang Fan, Sheng Zhang, Jingjing Zhang, Enrui Liu, Xinhui Li, Gangming Zhao, Zhao Lv
摘要
Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in the precise reconstruction of minute-level continuous imagined speech: traditional models struggle to balance the efficiency of temporal dependency modeling and information retention in long-sequence decoding. To address this issue, this paper proposes the Dynamic Multiscale Fusion Network (DMF2Mel), which consists of four core components: the Dynamic Contrastive Feature Aggregation Module (DC-FAM), the Hierarchical Attention-Guided Multi-Scale Network (HAMS-Net), the SplineMap attention mechanism, and the bidirectional state space module (convMamba). Specifically, the DC-FAM separates speech-related ''foreground features'' from noisy ''background features'' through local convolution and global attention mechanisms, effectively suppressing interference and enhancing the representation of transient signals. HAMS-Net, based on the U-Net framework, achieves cross-scale fusion of high-level semantics and low-level details. The SplineMap attention mechanism integrates the Adaptive Gated Kolmogorov-Arnold Network (AGKAN) to combine global context modeling with spline-based local fitting. The convMamba captures long-range temporal dependencies with linear complexity and enhances nonlinear dynamic modeling capabilities. Results on the SparrKULee dataset show that DMF2Mel achieves a Pearson correlation coefficient of 0.074 in mel spectrogram reconstruction for known subjects (a 48% improvement over the baseline) and 0.048 for unknown subjects (a 35% improvement over the baseline).Code is available at: https://github.com/fchest/DMF2Mel.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Open Vocabulary Electroencephalography-to-Text Decoding and Zero-Shot Sentiment ClassificationZhenhailong Wang, Heng JiAAAI 2022 · 被引用 122 次
- Towards Voice Reconstruction from EEG during Imagined SpeechYoung-Eun Lee, Seo-Hyun Lee, Sang-Ho Kim, Seong-Whan LeeAAAI 2023 · 被引用 63 次
- DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention DetectionSheng Yan, Cunhang Fan, Hongyu Zhang, Xiaoke Yang 等NeurIPS 2024 · 被引用 57 次
- ConDSeg: A General Medical Image Segmentation Framework via Contrast-Driven Feature EnhancementMengqi Lei, Haochen Wu, Xinhua Lv, Xin WangAAAI 2025 · 被引用 28 次
相关 Paper
- Entangled No More: Multi-Domain Decoupling for Robust Dynamic Graph Neural NetworksYouda Mo, Chaobo He, Junwei Cheng, Peng Mei 等ICML 2026
- MSFNet: Multi-Scale Fusion Network for Brain-Controlled Speaker ExtractionCunhang Fan, Jingjing Zhang, Hongyu Zhang, Wang Xiang 等ACM MM 2024 · 被引用 16 次
- IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationKai Li, Runxuan Yang, Fuchun Sun, Xiaolin HuICML 2024 · 被引用 28 次
- DHGCN: Dual HyperGraph Convolutional Network for EEG-Based Auditory Attention DetectionJian Zhou, Yingjie Xie, Cunhang Fan, Huabin Wang 等ACM MM 2025 · 被引用 5 次
- Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion ModelsGanxi Xu, Zhao-Rong Lai, Yuting Tang, Yonghao Song 等ICML 2026 · 被引用 1 次
