MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-Training
Dulhan Jayalath, ʻŌiwi Parker Jones
Abstract
Clinical brain-to-text interfaces are designed for paralysed patients who cannot provide extensive training recordings. Pre-training improves data-efficient generalisation by learning statistical priors across subjects, but these priors critically depend on context. While natural speech might unfold gradually over minutes, most methods pre-train with only a few seconds of context. Thus, we propose MEG-XL , a model pre-trained with 2.5 minutes of MEG context per sample, 5-300× longer than prior work, and equivalent to 191k tokens, capturing extended neural context. Fine-tuning on the task of word decoding from brain data, MEG-XL matches supervised performance with a fraction of the data (e.g. 1hr vs 50hrs) and outperforms brain foundation models. We find that models pre-trained with longer contexts learn representations that transfer better to word decoding. Our results indicate that long-context pre-training helps exploit extended neural context that other methods unnecessarily discard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ace1af20-9f0a-49b4-a36f-7fed89e60022Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- BIOT: Biosignal Transformer for Cross-data Learning in the WildChaoqi Yang, M. Brandon Westover, Jimeng SunNeurIPS 2023 · 345 citations
Related papers
- Brain-Inspired fMRI-to-Text Decoding via Incremental and Wrap-Up Language ModelingWentao Lu, Dong Nie, Pengcheng Xue, Zheng Cui et al.NeurIPS 2025 · 3 citations
- NeuroLM: A Universal Multi-task Foundation Model for Bridging the Gap between Language and EEG SignalsWeibang Jiang, Yansen Wang, Bao-Liang Lu, Dongsheng LiICLR 2025
- The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised LearningDulhan Jayalath, Gilad Landau, Brendan Shillingford, Mark W. Woolrich et al.ICML 2025
- Pretraining Large Brain Language Model for Active BCI: Silent SpeechJinzhao Zhou, Zehong Cao, Yiqun Duan, Connor Barkley et al.ACM MM 2025 · 2 citations
- A cross-species neural foundation model for end-to-end speech decodingYizi Zhang, Linyang He, Chaofei Fan, Tingkai Liu et al.ICLR 2026 · 5 citations
