One Hundred Neural Networks and Brains Watching Videos: Lessons from Alignment
Christina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris I. A. Groen
摘要
What can we learn from comparing video models to human brains, arguably the most efficient and effective video processing systems in existence? Our work takes a step towards answering this question by performing the first large-scale benchmarking of deep video models on representational alignment to the human brain, using publicly available models and a recently released video brain imaging (fMRI) dataset. We disentangle four factors of variation in the models (temporal modeling, classification task, architecture, and training dataset) that affect alignment to the brain, which we measure by conducting Representational Similarity Analysis across multiple brain regions and model layers. We show that temporal modeling is key for alignment to brain regions involved in early visual processing, while a relevant classification task is key for alignment to higher-level regions. Moreover, we identify clear differences between the brain scoring patterns across layers of CNNs and Transformers, and reveal how training dataset biases transfer to alignment with functionally selective brain areas. Additionally, we uncover a negative correlation of computational complexity to brain alignment. Measuring a total of 99 neural networks and 10 human brains watching videos, we aim to forge a path that widens our understanding of temporal and semantic video representations in brains and machines, ideally leading towards more efficient video models and more mechanistic explanations of processing in the human brain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative ModelsLarissa Höfling, Matthias Tangemann, Lotta Piefke, Susanne Keller 等ICLR 2026 · 被引用 5 次
- The Human Brain as a Dynamic Mixture of Expert Models in Video UnderstandingChristina Sartzetaki, Anne Zonneveld, Pablo Oyarzo, Alessandro T. Gifford 等ICLR 2026 · 被引用 4 次
- Quantifying Task-relevant Similarities in Representations Using Decision Variable CorrelationsYu Qian, Wilson S. Geisler, Xue-Xin WeiNeurIPS 2025 · 被引用 1 次
- Multimodal Scaling Laws for Task & Data-Optimized Models of Visual CortexAbdülkadir Gökce, Yingtian Tang, Martin SchrimpfICML 2026
它引用的顶会 Paper12
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos 等ICML 2021 · 被引用 1,021 次
相关 Paper
- Dimensionality Mismatch Between Brains and Artificial Neural NetworksSantiago Galella, Maren H. Wehrheim, Matthias KaschubeNeurIPS 2025
- Multi-modal brain encoding models for multi-modal stimuliSubba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Maneesh Kumar Singh 等ICLR 2025
- Disentangling the Factors of Convergence between Brains and DINOv3Joséphine Raugel, Marc Szafraniec, Huy V. Vo, Camille Couprie 等ICLR 2026
- SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching ExperimentsSimon Dahan, Gabriel Bénédict, Logan Zane John Williams, Yourong Guo 等ICLR 2025
- Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual CortexColin Conwell, David Mayo, Andrei Barbu, Michael A. Buice 等NeurIPS 2021 · 被引用 31 次
