Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals
Hui Zheng, Haiteng Wang, Wei-Bang Jiang, Zhongtao Chen, Li He, Pei-Yang Lin, Peng-Hu Wei, Guo-Guang Zhao, Yun-Zhe Liu
Abstract
Invasive brain-computer interfaces with Electrocorticography (ECoG) have shown promise for high-performance speech decoding in medical applications, but less damaging methods like intracranial stereo-electroencephalography (sEEG) remain underexplored. With rapid advances in representation learning, leveraging abundant recordings to enhance speech decoding is increasingly attractive. However, popular methods often pre-train temporal models based on brain-level tokens, overlooking that brain activities in different regions are highly desynchronized during tasks. Alternatively, they pre-train spatial-temporal models based on channel-level tokens but fail to evaluate them on challenging tasks like speech decoding, which requires intricate processing in specific language-related areas. To address this issue, we collected a well-annotated Chinese word-reading sEEG dataset targeting language-related brain networks from 12 subjects. Using this benchmark, we developed the Du-IN model, which extracts contextual embeddings based on region-level tokens through discrete codex-guided mask modeling. Our model achieves state-of-the-art performance on the 61-word classification task, surpassing all baselines. Model comparisons and ablation studies reveal that our design choices, including (i) temporal modeling based on region-level tokens by utilizing 1D depthwise convolution to fuse channels in the ventral sensorimotor cortex (vSMC) and superior temporal gyrus (STG) and (ii) self-supervision through discrete codex-guided mask modeling, significantly contribute to this performance. Overall, our approach -- inspired by neuroscience findings and capitalizing on region-level representations from specific brain regions -- is suitable for invasive brain modeling and represents a promising neuro-inspired AI approach in brain-computer interfaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2adf48f2-b740-4aed-a9ae-4b8eb5f37fa0Cited by top-tier papers8
- BaRISTA: Brain Scale Informed Spatiotemporal Representation of Human Intracranial Neural ActivityLucine L. Oganesian, Saba Hashemi, Maryam M. ShanechiNeurIPS 2025 · 7 citations
- EEGMirror: Leveraging EEG Data in the Wild Via Montage-Agnostic Self-Supervision for EEG to Video DecodingXuan-Hao Liu, Bao-Liang Lu, Wei-Long ZhengICCV 2025 · 5 citations
- Assembling the Mind's Mosaic: Towards EEG Semantic Intent DecodingJiahe Li, Junru Chen, Fanqi Shen, Jialan Yang et al.ICLR 2026 · 4 citations
- Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal TokenizationMohammad Hosseini, Eray Erturk, Saba Hashemi, Maryam ShanechiICML 2026 · 1 citation
- S³: Spiking Neurons as an Isolating Segmenter for Brain Signal DecodingQian Zheng, Ming Chen, Sha Zhao, Shi Gu et al.AAAI 2026
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
Related papers
- Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial RecordingsDi Wu, Siyuan Li, Chen Feng, Lu Cao et al.ICLR 2025
- CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG DecodingYuchen Zhou, Jiamin Wu, Zichen Ren, Zhouheng Yao et al.NeurIPS 2025 · 71 citations
- Pretraining Large Brain Language Model for Active BCI: Silent SpeechJinzhao Zhou, Zehong Cao, Yiqun Duan, Connor Barkley et al.ACM MM 2025 · 2 citations
- CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone DecodingYifan Zhuang, Calvin Huang, Zepeng Yu, Yongjie Zou et al.AAAI 2026
- NeuroCLUS: A Foundation Model with Functional Clustering for Intracranial Neural DecodingHui Zheng, Haiteng WangICML 2026
