Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech
Aditya R. Vaidya, Shailee Jain, Alexander Huth
Abstract
Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either hand-constructed acoustic filters or representations from supervised audio neural networks. In this work, we capi-talize on the progress of self-supervised speech representation learning (SSL) to create new state-of-the-art models of the human auditory system. Compared against acoustic baselines, phonemic features, and supervised models, representations from the middle layers of self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) con-sistently yield the best prediction performance for fMRI recordings within the auditory cortex (AC). Brain areas involved in low-level auditory processing exhibit a preference for earlier SSL model layers, whereas higher-level semantic areas prefer later layers. We show that these trends are due to the models’ ability to encode information at multiple linguistic levels (acoustic, phonetic, and lexical) along their representation depth. Overall, these results show that self-supervised models effectively capture the hierarchy of information relevant to different stages of speech processing in human cortex.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13f4e6aa-5c9c-4384-82a9-52f609920a1eCited by top-tier papers14
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec et al.NeurIPS 2022 · 164 citations
- Scaling laws for language encoding models in fMRIRichard J. Antonello, Aditya R. Vaidya, Alexander HuthNeurIPS 2023 · 137 citations
- Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech ModelsOmer Moussa, Mariya TonevaNeurIPS 2025 · 7 citations
- Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual DisentanglementLinyang He, Tianjun Zhong, Richard J. Antonello, Gavin Mischler et al.NeurIPS 2025 · 6 citations
- Scaling and context steer LLMs along the same computational path as the human brainJoséphine Raugel, Jérémy Rapin, Stéphane d'Ascoli, Valentin Wyart et al.NeurIPS 2025 · 6 citations
Builds on2
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechShailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel et al.NeurIPS 2020 · 58 citations
Related papers
- Do self-supervised speech models develop human-like perception biases?Juliette Millet, Ewan DunbarACL 2022 · 27 citations
- Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech RepresentationsWeiwei Lin, Chenhang He, Man-Wai Mak, Youzhi TuICML 2023 · 6 citations
- Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit PredictionJiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov et al.ICLR 2024 · 39 citations
- From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation modelMarvin Lavechin, Thomas HueberEMNLP 2025
- Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech RepresentationsLinyang He, Qiaolin Wang, Xilin Jiang, Nima MesgaraniEMNLP 2025 · 1 citation
