Transformer brain encoders explain human high-level visual responses
Hossein Adeli, Minni Sun, Nikolaus Kriegeskorte
Abstract
A major goal of neuroscience is to understand brain computations during visual processing in naturalistic settings. A dominant approach is to use image-computable deep neural networks trained with different task objectives as a basis for linear encoding models. However, in addition to requiring estimation of a large number of linear encoding parameters, this approach ignores the structure of the feature maps both in the brain and the models. Recently proposed alternatives factor the linear mapping into separate sets of spatial and feature weights, thus finding static receptive fields for units, which is appropriate only for early visual areas. In this work, we employ the attention mechanism used in the transformer architecture to study how retinotopic visual features can be dynamically routed to category-selective areas in high-level visual processing. We show that this computational motif is significantly more powerful than alternative methods in predicting brain activity during natural scene viewing, across different feature basis models and modalities. We also show that this approach is inherently more interpretable as the attention-routing signals for different high-level categorical areas can be easily visualized for any input image. Given its high performance at predicting brain responses to novel images, the model deserves consideration as a candidate mechanistic model of how visual information from retinotopic maps is routed in the human brain based on the relevance of the input content to different category-selective regions. Our code is available at https://github.com/Hosseinadeli/transformer_brain_encoder/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad822a64-adf8-48eb-a49b-5338010acfb7Cited by top-tier papers5
- In Silico Mapping of Visual Categorical Selectivity Across the Whole BrainEthan Hwang, Hossein Adeli, Wenxuan Guo, Andrew F. Luo et al.NeurIPS 2025 · 7 citations
- Towards Interpretable Visual Decoding with Attention to Brain RepresentationsPinyuan Feng, Hossein Adeli, Wenxuan Guo, Fan Cheng et al.ICLR 2026 · 1 citation
- Beyond Grid-Locked Voxels: Neural Response Functions for Continuous Brain EncodingHaomiao Chen, Keith W Jamison, Mert R. Sabuncu, Amy KuceyeskiICLR 2026 · 1 citation
- Multimodal Scaling Laws for Task & Data-Optimized Models of Visual CortexAbdülkadir Gökce, Yingtian Tang, Martin SchrimpfICML 2026
- Generating metamers of human scene understandingRitik Raina, Abe Leite, Alexandros Graikos, Seoyoung Ahn et al.ICLR 2026
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Unifying Top-Down and Bottom-Up Scanpath Prediction Using TransformersZhibo Yang, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue et al.CVPR 2024
- Relating transformers to models and neural representations of the hippocampal formationJames C. R. Whittington, Joseph Warren, Tim E. J. BehrensICLR 2022 · 110 citations
- Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision TransformersAndrew F. Luo, Jacob Yeung, Rushikesh Zawar, Shaurya Dewan et al.ICLR 2025
- Transformer Interpretability Beyond Attention VisualizationHila Chefer, Shir Gur, Lior WolfCVPR 2021
- Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual CortexColin Conwell, David Mayo, Andrei Barbu, Michael A. Buice et al.NeurIPS 2021 · 31 citations
