Modeling dynamic social vision highlights gaps between deep learning and humans
Kathy Garcia, Emalie McMahon, Colin Conwell, Michael F. Bonner, Leyla Isik
摘要
Deep learning models trained on computer vision tasks are widely considered the most successful models of human vision to date. The majority of work that supports this idea evaluates how accurately these models predict brain and behavioral responses to static images of objects and natural scenes. Real-world vision, however, is highly dynamic, and far less work has focused on evaluating the accuracy of deep learning models in predicting responses to stimuli that move, and that involve more complicated, higher-order phenomena like social interactions. Here, we present a dataset of natural videos and captions involving complex multi-agent interactions, and we benchmark 350+ image, video, and language models on behavioral and neural responses to the videos. As with prior work, we find that many vision models reach the noise ceiling in predicting visual scene features and responses along the ventral visual stream (often considered the primary neural substrate of object and scene recognition). In contrast, image models poorly predict human action and social interaction ratings and neural responses in the lateral stream (a neural pathway increasingly theorized as specializing in dynamic, social vision). Language models (given human sentence captions of the videos) predict action and social ratings better than either image or video models, but they still perform poorly at predicting neural responses in the lateral stream. Together these results identify a major gap in AI's ability to match human social vision and highlight the importance of studying vision in dynamic, natural contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Human Brain as a Dynamic Mixture of Expert Models in Video UnderstandingChristina Sartzetaki, Anne Zonneveld, Pablo Oyarzo, Alessandro T. Gifford 等ICLR 2026 · 被引用 4 次
- One Hundred Neural Networks and Brains Watching Videos: Lessons from AlignmentChristina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris I. A. GroenICLR 2025
它引用的顶会 Paper1
相关 Paper
- MIMIC-Bench: Exploring the User-Like Thinking and Mimicking Capabilities of Multimodal Large Language ModelsJiajie Teng, Huiyu Duan, Sijing Wu, Jiarui Wang 等ICLR 2026
- Sparse components distinguish visual pathways & their alignment to neural networksAmmar I Marvi, Nancy Kanwisher, Meenakshi KhoslaICLR 2025
- PHASE: PHysically-grounded Abstract Social Events for Machine Social PerceptionAviv Netanyahu, Tianmin Shu, Boris Katz, Andrei Barbu 等AAAI 2021 · 被引用 44 次
- CATER: A diagnostic dataset for Compositional Actions & TEmporal ReasoningRohit Girdhar, Deva RamananICLR 2020 · 被引用 198 次
- LaVCa: LLM-assisted Visual Cortex CaptioningTakuya Matsuyama, Shinji Nishimoto, Yu TakagiICLR 2026 · 被引用 8 次
