DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
Sheng Yan, Cunhang Fan, Hongyu Zhang, Xiaoke Yang, Jianhua Tao, Zhao Lv
Abstract
At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current AAD algorithms overlook the spatial distribution information within EEG signals and lack the ability to capture long-range latent dependencies, limiting the model's ability to decode brain activity. To address these issues, this paper proposes a dual attention refinement network with spatiotemporal construction for AAD, named DARNet, which consists of the spatiotemporal construction module, dual attention refinement module, and feature fusion &classifier module. Specifically, the spatiotemporal construction module aims to construct more expressive spatiotemporal feature representations, by capturing the spatial distribution characteristics of EEG signals. The dual attention refinement module aims to extract different levels of temporal patterns in EEG signals and enhance the model's ability to capture long-range latent dependencies. The feature fusion &classifier module aims to aggregate temporal patterns and dependencies from different levels and obtain the final classification results. The experimental results indicate that compared to the state-of-the-art models, DARNet achieves an average classification accuracy improvement of 5.9% for 0.1s, 4.6% for 1s, and 3.9% for 2s on the DTU dataset. While maintaining excellent classification performance, DARNet significantly reduces the number of required parameters. Compared to the state-of-the-art models, DARNet reduces the parameter count by 91%. Code is available at: https://github.com/fchest/DARNet.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc0b99c6-13c0-4265-a212-697bbfcd5c9eCited by top-tier papers4
- DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram ReconstructionCunhang Fan, Sheng Zhang, Jingjing Zhang, Enrui Liu et al.ACM MM 2025 · 4 citations
- SM-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention DetectionJiaqi Wang, Zhengyu Ma, Xiongri Shen, Chenlin Zhou et al.NeurIPS 2025 · 3 citations
- PCRNet: Phase-aware Complex Refinement Network for EEG-based Auditory Attention DecodingXiran Chen, Xiaoke Yang, Jian Zhou, Zhao Lv et al.ICML 2026
- MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic AlignmentRUI LIU, Zhige Chen, Pengshu, Wenlong You et al.ICLR 2026
Builds on1
Related papers
- DHGCN: Dual HyperGraph Convolutional Network for EEG-Based Auditory Attention DetectionJian Zhou, Yingjie Xie, Cunhang Fan, Huabin Wang et al.ACM MM 2025 · 5 citations
- ASTDF-Net: Attention-Based Spatial-Temporal Dual-Stream Fusion Network for EEG-Based Emotion RecognitionPeiliang Gong, Ziyu Jia, Pengpai Wang, Yueying Zhou et al.ACM MM 2023 · 38 citations
- MSFNet: Multi-Scale Fusion Network for Brain-Controlled Speaker ExtractionCunhang Fan, Jingjing Zhang, Hongyu Zhang, Wang Xiang et al.ACM MM 2024 · 16 citations
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian et al.ACM MM 2021 · 154 citations
- LoCoNet: Long-Short Context Network for Active Speaker DetectionXizi Wang, Feng Cheng, Gedas BertasiusCVPR 2024 · 25 citations
