Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals
Hui Zheng, Haiteng Wang, Wei-Bang Jiang, Zhongtao Chen, Li He, Pei-Yang Lin, Peng-Hu Wei, Guo-Guang Zhao, Yun-Zhe Liu
摘要
Invasive brain-computer interfaces with Electrocorticography (ECoG) have shown promise for high-performance speech decoding in medical applications, but less damaging methods like intracranial stereo-electroencephalography (sEEG) remain underexplored. With rapid advances in representation learning, leveraging abundant recordings to enhance speech decoding is increasingly attractive. However, popular methods often pre-train temporal models based on brain-level tokens, overlooking that brain activities in different regions are highly desynchronized during tasks. Alternatively, they pre-train spatial-temporal models based on channel-level tokens but fail to evaluate them on challenging tasks like speech decoding, which requires intricate processing in specific language-related areas. To address this issue, we collected a well-annotated Chinese word-reading sEEG dataset targeting language-related brain networks from 12 subjects. Using this benchmark, we developed the Du-IN model, which extracts contextual embeddings based on region-level tokens through discrete codex-guided mask modeling. Our model achieves state-of-the-art performance on the 61-word classification task, surpassing all baselines. Model comparisons and ablation studies reveal that our design choices, including (i) temporal modeling based on region-level tokens by utilizing 1D depthwise convolution to fuse channels in the ventral sensorimotor cortex (vSMC) and superior temporal gyrus (STG) and (ii) self-supervision through discrete codex-guided mask modeling, significantly contribute to this performance. Overall, our approach -- inspired by neuroscience findings and capitalizing on region-level representations from specific brain regions -- is suitable for invasive brain modeling and represents a promising neuro-inspired AI approach in brain-computer interfaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- BaRISTA: Brain Scale Informed Spatiotemporal Representation of Human Intracranial Neural ActivityLucine L. Oganesian, Saba Hashemi, Maryam M. ShanechiNeurIPS 2025 · 被引用 7 次
- EEGMirror: Leveraging EEG Data in the Wild Via Montage-Agnostic Self-Supervision for EEG to Video DecodingXuan-Hao Liu, Bao-Liang Lu, Wei-Long ZhengICCV 2025 · 被引用 5 次
- Assembling the Mind's Mosaic: Towards EEG Semantic Intent DecodingJiahe Li, Junru Chen, Fanqi Shen, Jialan Yang 等ICLR 2026 · 被引用 4 次
- Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal TokenizationMohammad Hosseini, Eray Erturk, Saba Hashemi, Maryam ShanechiICML 2026 · 被引用 1 次
- S³: Spiking Neurons as an Isolating Segmenter for Brain Signal DecodingQian Zheng, Ming Chen, Sha Zhao, Shi Gu 等AAAI 2026
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
相关 Paper
- Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial RecordingsDi Wu, Siyuan Li, Chen Feng, Lu Cao 等ICLR 2025
- CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG DecodingYuchen Zhou, Jiamin Wu, Zichen Ren, Zhouheng Yao 等NeurIPS 2025 · 被引用 71 次
- Pretraining Large Brain Language Model for Active BCI: Silent SpeechJinzhao Zhou, Zehong Cao, Yiqun Duan, Connor Barkley 等ACM MM 2025 · 被引用 2 次
- CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone DecodingYifan Zhuang, Calvin Huang, Zepeng Yu, Yongjie Zou 等AAAI 2026
- NeuroCLUS: A Foundation Model with Functional Clustering for Intracranial Neural DecodingHui Zheng, Haiteng WangICML 2026
