Disentangling syntax and semantics in the brain with deep networks
Charlotte Caucheteux, Alexandre Gramfort, Jean-Remi King
Abstract
The activations of language transformers like GPT-2 have been shown to linearly map onto brain activity during speech comprehension. However, the nature of these activations remains largely unknown and presumably conflate distinct linguistic classes. Here, we propose a taxonomy to factorize the high-dimensional activations of language models into four combinatorial classes: lexical, compositional, syntactic, and semantic representations. We then introduce a statistical method to decompose, through the lens of GPT-2's activations, the brain activity of 345 subjects recorded with functional magnetic resonance imaging (fMRI) during the listening of 4.6 hours of narrated text. The results highlight two findings. First, compositional representations recruit a more widespread cortical network than lexical ones, and encompass the bilateral temporal, parietal and prefrontal cortices. Second, contrary to previous claims, syntax and semantics are not associated with separated modules, but, instead, appear to share a common and distributed neural substrate. Overall, this study introduces a versatile framework to isolate, in the brain activity, the distributed representations of linguistic constructs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed8e4a62-b641-46c7-8e9b-3f9e6998cffbCited by top-tier papers13
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec et al.NeurIPS 2022 · 164 citations
- Can fMRI reveal the representation of syntactic structure in the brain?Aniketh Janardhan Reddy, Leila WehbeNeurIPS 2021 · 54 citations
- Neural Language Models are not Born Equal to Fit Brain Data, but Training HelpsAlexandre Pasquiou, Yair Lakretz, John T. Hale, Bertrand Thirion et al.ICML 2022 · 44 citations
- A Polar coordinate system represents syntax in large language modelsPablo Diego-Simón, Stéphane d'Ascoli, Emmanuel Chemla, Yair Lakretz et al.NeurIPS 2024 · 27 citations
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
Builds on2
Related papers
- Coupling Artificial Neurons in BERT and Biological Neurons in the Human BrainXu Liu, Mengyue Zhou, Gaosheng Shi, Yu Du et al.AAAI 2023 · 18 citations
- Unveiling Multi-level and Multi-modal Semantic Representations in the Human Brain using Large Language ModelsYuko Nakagi, Takuya Matsuyama, Naoko Koide-Majima, Hiroto Yamaguchi et al.EMNLP 2024 · 7 citations
- Probing Brain Activation Patterns by Dissociating Semantics and Syntax in SentencesShaonan Wang, Jiajun Zhang, Nan Lin, Chengqing ZongAAAI 2020 · 23 citations
- Do Large Language Models Think like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRIYu Lei, Xingyang Ge, Yi Zhang, Yiming Yang et al.AAAI 2026 · 2 citations
- fMRI predictors based on language models of increasing complexity recover brain left lateralizationLaurent Bonnasse-Gahot, Christophe PallierNeurIPS 2024 · 15 citations
