From Language to Cognition: How LLMs Outgrow the Human Language Network
Badr AlKhamissi, Greta Tuckute, Yingtian Tang, Taha Osama A Binhuraib, Antoine Bosselut, Martin Schrimpf
Abstract
Large language models (LLMs) exhibit remarkable similarity to neural activity in the human language network. However, the key properties of language underlying this alignmentand how brain-like representations emerge and change across training-remain unclear. We here benchmark 34 training checkpoints spanning 300B tokens across 8 different model sizes to analyze how brain alignment relates to linguistic competence. Specifically, we find that brain alignment tracks the development of formal linguistic competence-i.e., knowledge of linguistic rules-more closely than functional linguistic competence. While functional competence, which involves world knowledge and reasoning, continues to develop throughout training, its relationship with brain alignment is weaker, suggesting that the human language network primarily encodes formal linguistic structure rather than broader cognitive functions. Notably, we find that the correlation between next-word prediction, behavioral alignment, and brain alignment fades once models surpass human language proficiency. We further show that model size is not a reliable predictor of brain alignment when controlling for the number of features. Finally, using the largest set of rigorous neural language benchmarks to date, we show that language brain alignment benchmarks remain unsaturated, highlighting opportunities for improving future models. Taken together, our findings suggest that the human language network is best modeled by formal, rather than functional, aspects of language. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like SpecializationBadr AlKhamissi, C. Nicolò De Sabbata, Greta Tuckute, Zeming Chen et al.ICLR 2026 · 12 citations
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI UnderstandingYuxiang Wei, Yanteng Zhang, Xi Xiao, Chengxuan Qian et al.CVPR 2026 · 11 citations
- Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM PretrainingDeniz Bayazit, Aaron Mueller, Antoine BosselutACL 2026 · 3 citations
- On Emergent Social World Models - Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language ModelsPolina Tsvilodub, Jan-Felix Klumpp, Amir Pour, Jennifer Hu et al.ACL 2026 · 1 citation
- Language Models Grow Less Humanlike beyond Phase TransitionTatsuya Aoyama, Ethan WilcoxACL 2025
Builds on8
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
- Joint processing of linguistic properties in brains and language modelsSubba Reddy Oota, Manish Gupta, Mariya TonevaNeurIPS 2023 · 64 citations
- Neural Language Models are not Born Equal to Fit Brain Data, but Training HelpsAlexandre Pasquiou, Yair Lakretz, John T. Hale, Bertrand Thirion et al.ICML 2022 · 44 citations
Related papers
- When Language Models Lose Their Mind: The Consequences of Brain MisalignmentGabriele Merlin, Mariya TonevaICLR 2026 · 3 citations
- Linguistic Properties and Model Scale in Brain Encoding: From Small to Compressed Language ModelsSubba Reddy Oota, Satya Sai Srinath Namburi GNVV, Vijay Rowtula, Khushbu Pahwa et al.ICML 2026
- Do Large Language Models Think like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRIYu Lei, Xingyang Ge, Yi Zhang, Yiming Yang et al.AAAI 2026 · 2 citations
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for MeaningChen Shani, Liron Soffer, Dan Jurafsky, Yann LeCun et al.ICLR 2026 · 38 citations
- Scaling and context steer LLMs along the same computational path as the human brainJoséphine Raugel, Jérémy Rapin, Stéphane d'Ascoli, Valentin Wyart et al.NeurIPS 2025 · 6 citations
