Lune

ICLR2026Top-tier venue

A cross-species neural foundation model for end-to-end speech decoding

Yizi Zhang, Linyang He, Chaofei Fan, Tingkai Liu, Han Yu, Trung Le, Jingyuan Li, Scott W. Linderman, Lea Duncker, Francis R. Willett, Nima Mesgarani, Liam Paninski

2026Year
5Citations

Abstract

Speech brain-computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most systems use cascaded frameworks that decode phonemes before assembling sentences with an n-gram language model (LM), preventing joint optimization of all stages simultaneously. Here, we introduce an end-to-end BraIn-to-Text (BIT) framework that translates neural activity into coherent sentences using a single differentiable neural network. Central to our approach is a cross-task, cross-species pretrained neural encoder, whose representations transfer to both attempted and imagined speech. In a cascaded setting with an n-gram LM, the pretrained encoder establishes a new state-of-the-art (SOTA) on the Brain-to-Text '24 and '25 benchmarks. Integrated end-to-end with audio large language models (LLMs) and trained with contrastive learning for cross-modal alignment, BIT reduces the word error rate (WER) of the prior end-to-end method from 24.69% to 10.22%. Notably, we find that small-scale audio-LLMs markedly improve end-to-end decoding. Beyond record-setting performance, BIT aligns attempted and imagined speech embeddings to enable cross-task generalization. Altogether, our approach advances the integration of large, diverse neural datasets, paving the way for an end-to-end decoding framework that supports seamless, differentiable optimization. Project code: github.com/yzhang511/bit Recently, Feng et al. (2024) connect RNNs with LLMs to translate neural activity into sentences in an end-to-end manner. While promising, this approach does not explore modern architectures, such as transformers, which may better capture complex neural representations and enhance decoding performance. In contrast, Feghhi et al. (2025) incorporate transformers to decode phonemes and apply time masking to mitigate overfitting, achieving higher performance, but their study does not

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5e0cdc0b-0934-40ef-968c-a52c3cd9c989

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines