Language models and brains align due to more than next-word prediction and word-level information
Gabriele Merlin, Mariya Toneva
Abstract
Pretrained language models have been shown to significantly predict brain recordings of people comprehending language. Recent work suggests that the prediction of the next word is a key mechanism that contributes to this alignment. What is not yet understood is whether prediction of the next word is necessary for this observed alignment or simply sufficient, and whether there are other shared mechanisms or information that are similarly important. In this work, we take a step towards understanding the reasons for brain alignment via two simple perturbations in popular pretrained language models. These perturbations help us design contrasts that can control for different types of information. By contrasting the brain alignment of these differently perturbed models, we show that improvements in alignment with brain recordings are due to more than improvements in next-word prediction and word-level information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech ModelsOmer Moussa, Mariya TonevaNeurIPS 2025 · 7 citations
- When Language Models Lose Their Mind: The Consequences of Brain MisalignmentGabriele Merlin, Mariya TonevaICLR 2026 · 3 citations
- Fine-grained Analysis of Brain-LLM Alignment through Input AttributionMichela Proietti, Roberto Capobianco, Mariya TonevaICML 2026
- Improving Semantic Understanding in Speech Language Models via Brain-tuningOmer Moussa, Dietrich Klakow, Mariya TonevaICLR 2025
Builds on5
- Joint processing of linguistic properties in brains and language modelsSubba Reddy Oota, Manish Gupta, Mariya TonevaNeurIPS 2023 · 64 citations
- Can fMRI reveal the representation of syntactic structure in the brain?Aniketh Janardhan Reddy, Leila WehbeNeurIPS 2021 · 54 citations
- Sorting through the noise: Testing robustness of information processing in pre-trained language modelsLalchand Pandia, Allyson EttingerEMNLP 2021 · 20 citations
- Training language models to summarize narratives improves brain alignmentKhai Loong Aw, Mariya TonevaICLR 2023 · 11 citations
- UnNatural Language InferenceKoustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, Adina WilliamsACL 2021
Related papers
- Linguistic Properties and Model Scale in Brain Encoding: From Small to Compressed Language ModelsSubba Reddy Oota, Satya Sai Srinath Namburi GNVV, Vijay Rowtula, Khushbu Pahwa et al.ICML 2026
- Speech language models lack important brain-relevant semanticsSubba Reddy Oota, Emin Çelik, Fatma Deniz, Mariya TonevaACL 2024
- Abstraction Induces the Brain Alignment of Language and Speech ModelsEmily Cheng, Aditya Vaidya, Richard AntonelloICML 2026
- Neural Language Models are not Born Equal to Fit Brain Data, but Training HelpsAlexandre Pasquiou, Yair Lakretz, John T. Hale, Bertrand Thirion et al.ICML 2022 · 44 citations
- Explaining How Transformers Use Context to Build PredictionsJavier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, Marta R. Costa-jussàACL 2023 · 9 citations
