Modeling dynamic social vision highlights gaps between deep learning and humans
Kathy Garcia, Emalie McMahon, Colin Conwell, Michael F. Bonner, Leyla Isik
Abstract
Deep learning models trained on computer vision tasks are widely considered the most successful models of human vision to date. The majority of work that supports this idea evaluates how accurately these models predict brain and behavioral responses to static images of objects and natural scenes. Real-world vision, however, is highly dynamic, and far less work has focused on evaluating the accuracy of deep learning models in predicting responses to stimuli that move, and that involve more complicated, higher-order phenomena like social interactions. Here, we present a dataset of natural videos and captions involving complex multi-agent interactions, and we benchmark 350+ image, video, and language models on behavioral and neural responses to the videos. As with prior work, we find that many vision models reach the noise ceiling in predicting visual scene features and responses along the ventral visual stream (often considered the primary neural substrate of object and scene recognition). In contrast, image models poorly predict human action and social interaction ratings and neural responses in the lateral stream (a neural pathway increasingly theorized as specializing in dynamic, social vision). Language models (given human sentence captions of the videos) predict action and social ratings better than either image or video models, but they still perform poorly at predicting neural responses in the lateral stream. Together these results identify a major gap in AI's ability to match human social vision and highlight the importance of studying vision in dynamic, natural contexts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e4ddea7-f83a-4d44-af11-5333a3dee6a2Cited by top-tier papers2
- The Human Brain as a Dynamic Mixture of Expert Models in Video UnderstandingChristina Sartzetaki, Anne Zonneveld, Pablo Oyarzo, Alessandro T. Gifford et al.ICLR 2026 · 4 citations
- One Hundred Neural Networks and Brains Watching Videos: Lessons from AlignmentChristina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris I. A. GroenICLR 2025
Builds on1
Related papers
- MIMIC-Bench: Exploring the User-Like Thinking and Mimicking Capabilities of Multimodal Large Language ModelsJiajie Teng, Huiyu Duan, Sijing Wu, Jiarui Wang et al.ICLR 2026
- Sparse components distinguish visual pathways & their alignment to neural networksAmmar I Marvi, Nancy Kanwisher, Meenakshi KhoslaICLR 2025
- PHASE: PHysically-grounded Abstract Social Events for Machine Social PerceptionAviv Netanyahu, Tianmin Shu, Boris Katz, Andrei Barbu et al.AAAI 2021 · 44 citations
- CATER: A diagnostic dataset for Compositional Actions & TEmporal ReasoningRohit Girdhar, Deva RamananICLR 2020 · 198 citations
- LaVCa: LLM-assisted Visual Cortex CaptioningTakuya Matsuyama, Shinji Nishimoto, Yu TakagiICLR 2026 · 8 citations
