Disentangling syntax and semantics in the brain with deep networks
Charlotte Caucheteux, Alexandre Gramfort, Jean-Remi King
摘要
The activations of language transformers like GPT-2 have been shown to linearly map onto brain activity during speech comprehension. However, the nature of these activations remains largely unknown and presumably conflate distinct linguistic classes. Here, we propose a taxonomy to factorize the high-dimensional activations of language models into four combinatorial classes: lexical, compositional, syntactic, and semantic representations. We then introduce a statistical method to decompose, through the lens of GPT-2's activations, the brain activity of 345 subjects recorded with functional magnetic resonance imaging (fMRI) during the listening of 4.6 hours of narrated text. The results highlight two findings. First, compositional representations recruit a more widespread cortical network than lexical ones, and encompass the bilateral temporal, parietal and prefrontal cortices. Second, contrary to previous claims, syntax and semantics are not associated with separated modules, but, instead, appear to share a common and distributed neural substrate. Overall, this study introduces a versatile framework to isolate, in the brain activity, the distributed representations of linguistic constructs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec 等NeurIPS 2022 · 被引用 164 次
- Can fMRI reveal the representation of syntactic structure in the brain?Aniketh Janardhan Reddy, Leila WehbeNeurIPS 2021 · 被引用 54 次
- Neural Language Models are not Born Equal to Fit Brain Data, but Training HelpsAlexandre Pasquiou, Yair Lakretz, John T. Hale, Bertrand Thirion 等ICML 2022 · 被引用 44 次
- A Polar coordinate system represents syntax in large language modelsPablo Diego-Simón, Stéphane d'Ascoli, Emmanuel Chemla, Yair Lakretz 等NeurIPS 2024 · 被引用 27 次
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello 等NeurIPS 2024 · 被引用 26 次
它引用的顶会 Paper2
相关 Paper
- Coupling Artificial Neurons in BERT and Biological Neurons in the Human BrainXu Liu, Mengyue Zhou, Gaosheng Shi, Yu Du 等AAAI 2023 · 被引用 18 次
- Unveiling Multi-level and Multi-modal Semantic Representations in the Human Brain using Large Language ModelsYuko Nakagi, Takuya Matsuyama, Naoko Koide-Majima, Hiroto Yamaguchi 等EMNLP 2024 · 被引用 7 次
- Probing Brain Activation Patterns by Dissociating Semantics and Syntax in SentencesShaonan Wang, Jiajun Zhang, Nan Lin, Chengqing ZongAAAI 2020 · 被引用 23 次
- Do Large Language Models Think like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRIYu Lei, Xingyang Ge, Yi Zhang, Yiming Yang 等AAAI 2026 · 被引用 2 次
- fMRI predictors based on language models of increasing complexity recover brain left lateralizationLaurent Bonnasse-Gahot, Christophe PallierNeurIPS 2024 · 被引用 15 次
