The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised Learning
Dulhan Jayalath, Gilad Landau, Brendan Shillingford, Mark W. Woolrich, Oiwi Parker Jones
Abstract
The past few years have produced a series of spectacular advances in the decoding of speech from brain activity. The engine of these advances has been the acquisition of labelled data, with increasingly large datasets acquired from single subjects. However, participants exhibit individual differences, such as anatomy, and datasets use varied scanners and task designs. As a result, prior work has struggled to leverage data from multiple subjects, multiple datasets, multiple tasks, and unlabelled datasets. In turn, the field has not benefited from the rapidly growing number of open neural data repositories to exploit large-scale data and deep learning. To address this, we develop an initial set of neuroscience-inspired self-supervised objectives, together with a neural architecture, for representation learning from heterogeneous and unlabelled neural recordings. Experimental results show that representations learned with these objectives scale with data, generalise across subjects, datasets, and tasks, and outperform learning using only labelled data. In addition, we set new benchmarks for two foundational speech decoding tasks. Taken together, these methods now unlock the potential for training speech decoding models with orders of magnitude more existing data. Labelled MEG (scarce) Unlabelled MEG (abundant)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15f6469b-2e59-4621-9beb-7981ec6bd2dbCited by top-tier papers3
- Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech ModelsOmer Moussa, Mariya TonevaNeurIPS 2025 · 7 citations
- Towards Brain Passage Retrieval: An Investigation of EEG Query RepresentationsNiall McGuire, Yashar MoshfeghiSIGIR 2025 · 5 citations
- MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-TrainingDulhan Jayalath, ʻŌiwi Parker JonesICML 2026
Builds on5
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIWei-Bang Jiang, Li-Ming Zhao, Bao-Liang LuICLR 2024 · 298 citations
- Open Vocabulary Electroencephalography-to-Text Decoding and Zero-Shot Sentiment ClassificationZhenhailong Wang, Heng JiAAAI 2022 · 122 citations
- BrainBERT: Self-supervised representation learning for intracranial recordingsChristopher Wang, Vighnesh Subramaniam, Adam Uri Yaari, Gabriel Kreiman et al.ICLR 2023 · 13 citations
Related papers
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec et al.NeurIPS 2022 · 164 citations
- Population Transformer: Learning Population-level Representations of Neural ActivityGeeling Chau, Christopher Wang, Sabera J. Talukder, Vighnesh Subramaniam et al.ICLR 2025
- Real-World Unsupervised Models Generalize to Predict Brain Responses to Out-of-Distribution StimuliChenggang Chen, Zhiyu Yang, Xiaoqin WangICML 2026
- Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial RecordingsDi Wu, Siyuan Li, Chen Feng, Lu Cao et al.ICLR 2025
- A Unified, Scalable Framework for Neural Population DecodingMehdi Azabou, Vinam Arora, Venkataramana Ganesh, Ximeng Mao et al.NeurIPS 2023 · 136 citations
