Neural Foundations of Mental Simulation: Future Prediction of Latent Representations on Dynamic Scenes
Aran Nayebi, Rishi Rajalingham, Mehrdad Jazayeri, Guangyu Robert Yang
摘要
Humans and animals have a rich and flexible understanding of the physical world, which enables them to infer the underlying dynamical trajectories of objects and events, plausible future states, and use that to plan and anticipate the consequences of actions. However, the neural mechanisms underlying these computations are unclear. We combine a goal-driven modeling approach with dense neurophysiological data and high-throughput human behavioral readouts that contain thousands of comparisons to directly impinge on this question. Specifically, we construct and evaluate several classes of sensory-cognitive networks to predict the future state of rich, ethologically-relevant environments, ranging from self-supervised end-to-end models with pixel-wise or object-slot objectives, to models that future predict in the latent space of purely static image-pretrained or dynamic video-pretrained foundation models. We find that “scale is not all you need”, and that many state-of-the-art machine learning models fail to perform well on our neural and behavioral benchmarks for future prediction. In fact, only one class of models matches these data well overall. We find that neural responses are currently best predicted by models trained to predict the future state of their environment in the latent space of pretrained foundation models optimized for dynamic scenes in a self-supervised manner. These models also approach the neurons’ ability to predict the environmental state variables that are visually hidden from view, despite not being explicitly trained to do so. Finally, we find that not all foundation model latents are equal. Notably, models that future predict in the latent space of video foundation models that are optimized to support a diverse range of egocentric sensorimotor tasks, reasonably match both human behavioral error patterns and neural dynamics across all environmental scenarios that we were able to test. Overall, these findings suggest that the neural mechanisms and behaviors of primate mental simulation have strong inductive biases associated with them, and are thus far most consistent with being optimized to future predict on reusable visual representations that are useful for Embodied AI more generally.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Self-supervised video pretraining yields robust and more human-aligned visual representationsNikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. HénaffNeurIPS 2023 · 被引用 27 次
- Recurrent neural network dynamical systems for biological visionWayne Soo, Aldo Battista, Puria Radmard, Xiao-Jing WangNeurIPS 2024 · 被引用 7 次
- Intrinsic Goals for Autonomous Agents: Model-Based Exploration in Virtual Zebrafish Predicts Ethological Behavior and Whole-Brain DynamicsReece Keller, Alyn Kirsch, Felix Pei, Xaq Pitkow 等NeurIPS 2025 · 被引用 5 次
- A Deep Learning Model of Mental Rotation Informed by Interactive VR ExperimentsRaymond Khazoum, Daniela Fernandes, Aleksandr Krylov, Qin Li 等ICML 2026 · 被引用 2 次
- Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingFan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen 等ICML 2026
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying 等ICML 2020 · 被引用 1,439 次
相关 Paper
- Your head is there to move you around: Goal-driven models of the primate dorsal pathwayPatrick J. Mineault, Shahab Bakhtiari, Blake A. Richards, Christopher C. PackNeurIPS 2021 · 被引用 61 次
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear 等ICML 2020 · 被引用 88 次
- Modeling dynamic social vision highlights gaps between deep learning and humansKathy Garcia, Emalie McMahon, Colin Conwell, Michael F. Bonner 等ICLR 2025 · 被引用 7 次
- PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and PlanningAngel Villar-Corrales, Sven BehnkeICML 2025
- Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems ApproachSteeven Janny, Hervé Poirier, Leonid Antsfeld, Guillaume Bono 等CVPR 2025
