The Human Brain as a Dynamic Mixture of Expert Models in Video Understanding
Christina Sartzetaki, Anne Zonneveld, Pablo Oyarzo, Alessandro T. Gifford, Radoslaw Martin Cichy, Pascal Mettes, Iris I. A. Groen
摘要
The human brain is the most efficient and versatile system for processing dynamic visual input. By comparing representations from deep video models to brain activity, we can gain insights into mechanistic solutions for effective video processing, important to better understand the brain and to build better models. Current works in model-brain alignment primarily focus on fMRI measurements, leaving open questions about fine-grained dynamic processing. Here, we introduce the first large-scale model benchmarking on alignment to dynamic electroencephalography (EEG) recordings of short natural videos. We analyze 100+ models across the axes of temporal integration, classification task, architecture, and pretraining, using our proposed Cross-Temporal Representational Similarity Analysis (CT-RSA) which matches the best time-unfolded model features to dynamically evolving brain responses, distilling 107 alignment scores. Our findings reveal novel insights on how continuous visual input is integrated in the brain, beyond the standard temporal processing hierarchy from low to high-level representations. After initial alignment to hierarchical static object processing, responses in posterior electrodes best align to mid-level temporally-integrative action features, showing high temporal correspondence to feature timings. In contrast, responses in frontal electrodes best align with high-level static action representations and show no temporal correspondence to the video. Additionally, temporally-integrating state-space models show superior alignment to intermediate posterior activity, in which self-supervised pretraining is also beneficial. We draw a metaphor to a dynamic mixture of expert models for the changing neural preference in tasks and temporal integration reflected in the alignment to different model types across time. We posit that a single best-aligned model would need such training and architecture as to allow combining and dynamically switching between these capacities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Towards robust vision by multi-task learning on monkey visual cortexShahd Safarani, Arne Nix, Konstantin Willeke, Santiago A. Cadena 等NeurIPS 2021 · 被引用 67 次
- TRIBE: TRImodal Brain Encoder for whole-brain fMRI response predictionStéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Hubert Banville 等ICLR 2026 · 被引用 28 次
- Cognitive Steering in Deep Neural Networks via Long-Range Modulatory Feedback ConnectionsTalia Konkle, George A. AlvarezNeurIPS 2023 · 被引用 22 次
- Chirality in Action: Time-Aware Video Representation Learning by Latent StraighteningPiyush Bagad, Andrew ZissermanNeurIPS 2025 · 被引用 14 次
相关 Paper
- One Hundred Neural Networks and Brains Watching Videos: Lessons from AlignmentChristina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris I. A. GroenICLR 2025
- EEG2Video: Towards Decoding Dynamic Visual Perception from EEG SignalsXuan-Hao Liu, Yan-Kai Liu, Yansen Wang, Kan Ren 等NeurIPS 2024 · 被引用 59 次
- Long-Range Feedback Spiking Network Captures Dynamic and Static Representations of the Visual Cortex under Movie StimuliLiwei Huang, Zhengyu Ma, Liutao Yu, Huihui Zhou 等NeurIPS 2024 · 被引用 5 次
- CrossMind-VL: Multi-Subject Mind-to-Video Decoding with Multimodal LLM Semantic GroundingXuanliu Zhu, Yiqiao Chai, Runnan Li, Mingying Lan 等ACM MM 2025
- CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG DecodingYuchen Zhou, Jiamin Wu, Zichen Ren, Zhouheng Yao 等NeurIPS 2025 · 被引用 71 次
