Speech language models lack important brain-relevant semantics
Subba Reddy Oota, Emin Çelik, Fatma Deniz, Mariya Toneva
摘要
Despite known differences between reading and listening in the brain, recent work has shown that text-based language models predict both text-evoked and speech-evoked brain activity to an impressive degree. This poses the question of what types of information language models truly predict in the brain. We investigate this question via a direct approach, in which we systematically remove specific lowlevel stimulus features (textual, speech, and visual) from language model representations to assess their impact on alignment with fMRI brain recordings during reading and listening. Comparing these findings with speech-based language models reveals starkly different effects of low-level features on brain alignment. While text-based models show reduced alignment in early sensory regions post-removal, they retain significant predictive power in late language regions. In contrast, speech-based models maintain strong alignment in early auditory regions even after feature removal but lose all predictive power in late language regions. These results suggest that speech-based models provide insights into additional information processed by early auditory regions, but caution is needed when using them to model processing in late language regions. We make our code publicly available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Brain-Informed Fine-Tuning for Improved Multilingual Understanding in Language ModelsAnuja Negi, Subba Reddy Oota, Anwar Nunez-Elizalde, Manish Gupta 等NeurIPS 2025 · 被引用 9 次
- Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech ModelsOmer Moussa, Mariya TonevaNeurIPS 2025 · 被引用 7 次
- Meta-Learning an In-Context Transformer Model of Human Higher Visual CortexMuquan Yu, Mu Nan, Hossein Adeli, Jacob S. Prince 等NeurIPS 2025 · 被引用 5 次
- When Language Models Lose Their Mind: The Consequences of Brain MisalignmentGabriele Merlin, Mariya TonevaICLR 2026 · 被引用 3 次
- Meta-Learning In-Context Enables Training-Free Cross Subject Brain DecodingMu Nan, Muquan Yu, Weijian Mai, Jacob S. Prince 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper7
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec 等NeurIPS 2022 · 被引用 164 次
- Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to SpeechAditya R. Vaidya, Shailee Jain, Alexander HuthICML 2022 · 被引用 81 次
- Joint processing of linguistic properties in brains and language modelsSubba Reddy Oota, Manish Gupta, Mariya TonevaNeurIPS 2023 · 被引用 64 次
相关 Paper
- Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During ListeningPadakanti Srijith, Khushbu Pahwa, Radhika Mamidi, Bapi Raju Surampudi 等EMNLP 2025
- Language models and brains align due to more than next-word prediction and word-level informationGabriele Merlin, Mariya TonevaEMNLP 2024 · 被引用 2 次
- Improving Semantic Understanding in Speech Language Models via Brain-tuningOmer Moussa, Dietrich Klakow, Mariya TonevaICLR 2025
- Abstraction Induces the Brain Alignment of Language and Speech ModelsEmily Cheng, Aditya Vaidya, Richard AntonelloICML 2026
- fMRI predictors based on language models of increasing complexity recover brain left lateralizationLaurent Bonnasse-Gahot, Christophe PallierNeurIPS 2024 · 被引用 15 次
