Disentangling the Factors of Convergence between Brains and DINOv3
Joséphine Raugel, Marc Szafraniec, Huy V. Vo, Camille Couprie, Jérémy Rapin, Stéphane d'Ascoli, Patrick Labatut, Piotr Bojanowski, Valentin Wyart, Jean-Remi King
Abstract
Many AI models trained on natural images develop representations that resemble those of the human brain. However, the factors driving this brain-model similarity remain poorly understood. To disentangle how the model, training and data independently lead a neural network to develop brain-like representations, we train a family of self-supervised vision transformers (DINOv3) that systematically vary these factors. We compare their representations of images to those of the human brain recorded through fMRI and MEG, providing high resolution in both spatial and temporal analyses. We assess the brain-model similarity with three complementary metrics focusing on representational similarity, topographical organization, and temporal dynamics. We show that all three factors -model size, training amount, and image type -independently and interactively impact each of these brain similarity metrics. In particular, the largest DINOv3 models trained with the most human-centric images reach the highest brain-similarity. These findings generalize across seven additional models. This emergence of brain-like representations in AI models follows a specific chronology during training: models first align with the early representations of the sensory cortices, and only align with the late and prefrontal representations of the brain with considerably more training. Finally, this developmental trajectory is indexed by structural and functional properties of the human cortex: representations acquired last by the models specifically align with cortical areas with the largest developmental expansion, thickness, least myelination and slowest timescales. Overall, these findings disentangle the interplay between architecture and experience in shaping how artificial neural networks come to see the world as humans do, thus offering a promising framework to understand how the human brain comes to represent its visual world. Method Figure . A. We compare the activations of DINOv3 to the activations of the human brain in response to the same images. B. To understand the factors that steer DINOv3 towards brain activity, we train from scratch a variety of models on different image domains (human-centric, satellite or biological data), with varying amounts of data. C. We compare each model to both fMRI and MEG (high spatial and temporal resolutions) by computing the linear similarity of their representations and the similarity of their hierarchical organization (encoding, spatial and temporal scores).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4c36d81-be8f-4c1a-89fb-4587edc3c6daBuilds on5
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec et al.NeurIPS 2022 · 164 citations
- Brain encoding models based on multimodal transformers can transfer across language and visionJerry Tang, Meng Du, Vy A. Vo, Vasudev Lal et al.NeurIPS 2023 · 76 citations
- When Representations Align: Universality in Representation Learning DynamicsLoek van Rossem, Andrew M. SaxeICML 2024 · 8 citations
- Alignment between Brains and AI: Evidence for Convergent Evolution across Modalities, Scales and Training TrajectoriesGuobin Shen, Dongcheng Zhao, Yiting Dong, Qian Zhang et al.ICML 2026 · 6 citations
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov et al.CVPR 2022
Related papers
- One Hundred Neural Networks and Brains Watching Videos: Lessons from AlignmentChristina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris I. A. GroenICLR 2025
- Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual CortexColin Conwell, David Mayo, Andrei Barbu, Michael A. Buice et al.NeurIPS 2021 · 31 citations
- Brain decoding: toward real-time reconstruction of visual perceptionYohann Benchetrit, Hubert J. Banville, Jean-Remi KingICLR 2024 · 108 citations
- Dimensionality Mismatch Between Brains and Artificial Neural NetworksSantiago Galella, Maren H. Wehrheim, Matthias KaschubeNeurIPS 2025
- Scaling and context steer LLMs along the same computational path as the human brainJoséphine Raugel, Jérémy Rapin, Stéphane d'Ascoli, Valentin Wyart et al.NeurIPS 2025 · 6 citations
