A cross-species neural foundation model for end-to-end speech decoding
Yizi Zhang, Linyang He, Chaofei Fan, Tingkai Liu, Han Yu, Trung Le, Jingyuan Li, Scott W. Linderman, Lea Duncker, Francis R. Willett, Nima Mesgarani, Liam Paninski
Abstract
Speech brain-computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most systems use cascaded frameworks that decode phonemes before assembling sentences with an n-gram language model (LM), preventing joint optimization of all stages simultaneously. Here, we introduce an end-to-end BraIn-to-Text (BIT) framework that translates neural activity into coherent sentences using a single differentiable neural network. Central to our approach is a cross-task, cross-species pretrained neural encoder, whose representations transfer to both attempted and imagined speech. In a cascaded setting with an n-gram LM, the pretrained encoder establishes a new state-of-the-art (SOTA) on the Brain-to-Text '24 and '25 benchmarks. Integrated end-to-end with audio large language models (LLMs) and trained with contrastive learning for cross-modal alignment, BIT reduces the word error rate (WER) of the prior end-to-end method from 24.69% to 10.22%. Notably, we find that small-scale audio-LLMs markedly improve end-to-end decoding. Beyond record-setting performance, BIT aligns attempted and imagined speech embeddings to enable cross-task generalization. Altogether, our approach advances the integration of large, diverse neural datasets, paving the way for an end-to-end decoding framework that supports seamless, differentiable optimization. Project code: github.com/yzhang511/bit Recently, Feng et al. (2024) connect RNNs with LLMs to translate neural activity into sentences in an end-to-end manner. While promising, this approach does not explore modern architectures, such as transformers, which may better capture complex neural representations and enhance decoding performance. In contrast, Feghhi et al. (2025) incorporate transformers to decode phonemes and apply time masking to mitigate overfitting, achieving higher performance, but their study does not
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e0cdc0b-0934-40ef-968c-a52c3cd9c989Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- Enhancing EEG-to-Text Decoding through Transferable Representations from Pre-trained Contrastive EEG-Text Masked AutoencoderJiaqi Wang, Zhenxi Song, Zhengyu Ma, Xipeng Qiu et al.ACL 2024
- Assembling the Mind's Mosaic: Towards EEG Semantic Intent DecodingJiahe Li, Junru Chen, Fanqi Shen, Jialan Yang et al.ICLR 2026 · 4 citations
- MindLLM: A Subject-Agnostic and Versatile Model for fMRI-to-text DecodingWeikang Qiu, Zheng Huang, Haoyu Hu, Aosong Feng et al.ICML 2025
- Brain-Inspired fMRI-to-Text Decoding via Incremental and Wrap-Up Language ModelingWentao Lu, Dong Nie, Pengcheng Xue, Zheng Cui et al.NeurIPS 2025 · 3 citations
- MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-TrainingDulhan Jayalath, ʻŌiwi Parker JonesICML 2026
