Lune

ICLR2026Top-tier venue

Disentangling the Factors of Convergence between Brains and DINOv3

Joséphine Raugel, Marc Szafraniec, Huy V. Vo, Camille Couprie, Jérémy Rapin, Stéphane d'Ascoli, Patrick Labatut, Piotr Bojanowski, Valentin Wyart, Jean-Remi King

2026Year

Abstract

Many AI models trained on natural images develop representations that resemble those of the human brain. However, the factors driving this brain-model similarity remain poorly understood. To disentangle how the model, training and data independently lead a neural network to develop brain-like representations, we train a family of self-supervised vision transformers (DINOv3) that systematically vary these factors. We compare their representations of images to those of the human brain recorded through fMRI and MEG, providing high resolution in both spatial and temporal analyses. We assess the brain-model similarity with three complementary metrics focusing on representational similarity, topographical organization, and temporal dynamics. We show that all three factors -model size, training amount, and image type -independently and interactively impact each of these brain similarity metrics. In particular, the largest DINOv3 models trained with the most human-centric images reach the highest brain-similarity. These findings generalize across seven additional models. This emergence of brain-like representations in AI models follows a specific chronology during training: models first align with the early representations of the sensory cortices, and only align with the late and prefrontal representations of the brain with considerably more training. Finally, this developmental trajectory is indexed by structural and functional properties of the human cortex: representations acquired last by the models specifically align with cortical areas with the largest developmental expansion, thickness, least myelination and slowest timescales. Overall, these findings disentangle the interplay between architecture and experience in shaping how artificial neural networks come to see the world as humans do, thus offering a promising framework to understand how the human brain comes to represent its visual world. Method Figure . A. We compare the activations of DINOv3 to the activations of the human brain in response to the same images. B. To understand the factors that steer DINOv3 towards brain activity, we train from scratch a variety of models on different image domains (human-centric, satellite or biological data), with varying amounts of data. C. We compare each model to both fMRI and MEG (high spatial and temporal resolutions) by computing the linear similarity of their representations and the similarity of their hierarchical organization (encoding, spatial and temporal scores).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e4c36d81-be8f-4c1a-89fb-4587edc3c6da

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines