Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream
Abdülkadir Gökce, Martin Schrimpf
Abstract
When trained on large-scale object classification datasets, certain artificial neural network models begin to approximate core object recognition behaviors and neural response patterns in the primate brain. While recent machine learning advances suggest that scaling compute, model size, and dataset size improves task performance, the impact of scaling on brain alignment remains unclear. In this study, we explore scaling laws for modeling the primate visual ventral stream by systematically evaluating over 600 models trained under controlled conditions on benchmarks spanning V1, V2, V4, IT and behavior. We find that while behavioral alignment continues to scale with larger models, neural alignment saturates. This observation remains true across model architectures and training datasets, even though models with stronger inductive biases and datasets with higher-quality images are more computeefficient. Increased scaling is especially beneficial for higher-level visual areas, where small models trained on few samples exhibit only poor alignment. Our results suggest that while scaling current architectures and datasets might suffice for alignment with human core object recognition behavior, it will not yield improved models of the brain's visual ventral stream, highlighting the need for novel strategies in building brain models. The advent of neural networks has revolutionized our understanding and modeling of complex neural processes. A particularly active area of study is the ventral visual stream in primates, a key pathway in the brain responsible for processing visual information (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3207475c-c678-4215-a396-bda4bdab78e5Cited by top-tier papers6
- OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural TokensKonstantin Friedrich Willeke, Polina Turishcheva, Alex Gilbert, Goirik Chakrabarty et al.ICLR 2026 · 10 citations
- Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN RobustnessLucas Piper, Arlindo L. Oliveira, Tiago MarquesNeurIPS 2025 · 4 citations
- Inducing Dyslexia in Vision Language ModelsMelika Honarmand, Ayati Sharma, Badr AlKhamissi, Johannes Mehrer et al.ICLR 2026 · 3 citations
- Multimodal Scaling Laws for Task & Data-Optimized Models of Visual CortexAbdülkadir Gökce, Yingtian Tang, Martin SchrimpfICML 2026
- Model-Guided Microstimulation Steers Primate Visual BehaviorJohannes Mehrer, Ben Lonnqvist, Anna Mitola, Paolo Papale et al.ICLR 2026
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Dimensionality Mismatch Between Brains and Artificial Neural NetworksSantiago Galella, Maren H. Wehrheim, Matthias KaschubeNeurIPS 2025
- Human alignment of neural network representationsLukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen et al.ICLR 2023 · 15 citations
- Vision CNNs trained to estimate spatial latents learned similar ventral-stream-aligned representationsYudi Xie, Weichen Huang, Esther Alter, Jeremy Schwartz et al.ICLR 2025
- Wiring Up Vision: Minimizing Supervised Synaptic Updates Needed to Produce a Primate Ventral StreamFranziska Geiger, Martin Schrimpf, Tiago Marques, James J. DiCarloICLR 2022 · 14 citations
- Aligning Model and Macaque Inferior Temporal Cortex Representations Improves Model-to-Human Behavioral Alignment and Adversarial RobustnessJoel Dapello, Kohitij Kar, Martin Schrimpf, Robert Baldwin Geary et al.ICLR 2023 · 27 citations
