ICML2026
Alignment between Brains and AI: Evidence for Convergent Evolution across Modalities, Scales and Training Trajectories
Guobin Shen, Dongcheng Zhao, Yiting Dong, Qian Zhang, Yi Zeng
被引用 6 次
摘要
Artificial and biological systems may evolve similar computational solutions despite fundamental differences in architecture and learning mechanisms-a form of convergent evolution. We demonstrate this phenomenon through large-scale analysis of alignment between human brain activity and internal representations of over 600 AI models spanning language and vision domains, from 1.33M to 72B parameters. Analyzing 60 million alignment measurements reveals that higher-performing models spontaneously develop stronger brain alignment without explicit neural constraints, with language models showing markedly stronger correlation (r = 0.89, p < 7.5 × 10 -13 ) than vision models (r = 0.53, p < 2.0 × 10 -44 ). Crucially, longitudinal analysis demonstrates that brain alignment consistently precedes performance improvements during training, suggesting that developing brain-like representations may be a necessary stepping stone toward higher capabilities. We find systematic patterns: language models exhibit strongest alignment with limbic and integrative regions, while vision models show progressive alignment with visual cortices; deeper processing layers converge across modalities; and as representational scale increases, alignment systematically shifts from primary sensory to higher-order associative regions. These findings provide compelling evidence that optimization for task performance naturally drives AI systems toward brain-like computational strategies, offering both fundamental insights into principles of intelligent information processing and practical guidance for developing more capable AI systems. mechanisms underlying the emergence of brain-like computational strategies, providing insights into the fundamental nature of intelligence. Convergent evolution provides a powerful framework for understanding brain-AI alignment. Just as vertebrates and cephalopods independently evolved camera-like eyes under similar visual processing demands 14, 15 , artificial and biological systems may converge on similar computational strategies when facing shared information-processing challenges [16] [17] [18] . This convergence has profound implications: high-performing AI models offer testbeds for understanding biological computation 18, 19 , while brain organization may guide development of more robust and efficient AI systems 16 . Brain-AI alignment thus serves as both a diagnostic tool and a blueprint for future intelligence architectures. To systematically test this convergent evolution hypothesis and address the limitations of prior work, we present a comprehensive analysis of brain-AI alignment unprecedented in scale. Our study examines internal representations from over 600 models across diverse architectures, scales, training trajectories, and modalities. Through analyzing 60 million alignment measurements, we directly compare layer-wise activations to human neural recordings, addressing three fundamental questions: (1) Does brain alignment precede performance improvements during training, suggesting it may be a necessary stepping stone? (2) Do alignment patterns differ systematically between vision and language modalities, or converge toward universal principles? (3) How do model hierarchies progressively map onto cortical processing levels from sensory to associative regions? Our findings reveal that artificial and biological intelligence, despite their distinct evolutionary paths, indeed converge toward similar computational solutions. This establishes fundamental principles that could transform both our understanding of intelligence and the development of future AI systems. Results We developed a systematic large-scale framework to comprehensively assess brain-AI representational alignment across sensory modalities, model scales, and training dynamics (Figure 1 ). Our analysis leveraged fMRI recordings from the Natural Scenes Dataset (NSD) 20 , which captures neural activity from multiple subjects viewing thousands of naturalistic images. These images, sourced from the COCO dataset 21 and paired with human-generated captions, enabled multimodal analyses across both vision and language domains. Our AI model collection comprised 630 neural networks spanning diverse architectures and scales: 36 large language models (0.5-72B parameters) including Qwen 6 , Llama 5 , and Gemma 22 families, and 594 vision models (1.33-1014M parameters) from CNNs to transformers 7, 23 . To quantify alignment, we employed Centered Kernel Alignment (CKA) 24 across multiple spatial scales, mapping model representations to brain regions defined by the HCP_MMP1 parcellation 25 and Yeo-7 functional networks 26 . For longitudinal analyses, we tracked representational evolution in the Pythia language model family 27 and MixNet vision model family 28 throughout training. Correlation Patterns Between Brain Alignment and Model Effectiveness Our first key finding is a robust, positive correlation between model performance and brain alignment across