SVBench: Evaluation of Video Generation Models on Social Reasoning
Wenshuo Peng, Gongxuan Wang, Tianmeng Yang, Chuanhao Li, Xiaojie Xu, Hui He, Kaipeng Zhang
Abstract
Recent text-to-video generation models have made remarkable progress in visual realism, motion fidelity, and text-video alignment, yet they still struggle to produce socially coherent behavior. Unlike humans, who readily infer intentions, beliefs, emotions, and social norms from brief visual cues, current models often generate literal scenes without capturing the underlying causal and psychological dynamics. To systematically assess this limitation, we introduce the first benchmark for social reasoning in video generation. Grounded in developmental and social psychology, the benchmark covers thirty classic social cognition paradigms spanning seven core dimensions: mental-state inference, goal-directed action, joint attention, social coordination, prosocial behavior, social norms, and multi-agent strategy. To operationalize these paradigms, we build a fully training-free agent-based pipeline that distills the reasoning structure of each paradigm, synthesizes diverse video-ready scenarios, enforces conceptual neutrality and difficulty control through cue-based critique, and evaluates generated videos with a high-capacity VLM judge along five interpretable dimensions of social reasoning. Using this framework, we conduct the first large-scale evaluation of seven state-of-the-art video generation systems. Results show a clear gap between surface-level plausibility and deeper social reasoning, suggesting that current models remain limited in their ability to generate socially grounded behavior. https://github.com/Gloria2tt/SVBench-Evaluation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3903636-64c2-41b2-a09d-ec41286161e3Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Improving Video Generation with Human FeedbackJie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan et al.NeurIPS 2025 · 284 citations
- Evaluating Newtonian Mechanics in Video Generative Models with Real Physical SystemsAntonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii, Antonis Vozikis et al.ICML 2026 · 39 citations
- Data Adaptive Traceback for Vision-Language Foundation Models in Image ClassificationWenshuo Peng, Kaipeng Zhang, Yue Yang, Hao Zhang et al.AAAI 2024 · 4 citations
- OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language ModelsHainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du et al.ACL 2024
- VBench: Comprehensive Benchmark Suite for Video Generative ModelsZiqi Huang, Yinan He, Jiashuo Yu, Fan Zhang et al.CVPR 2024
Related papers
- TiViBench: Benchmarking Think-in-Video Reasoning for Video GenerationHarold Haodong Chen, Disen Lan, Wen-Jie Shu, Qingyang Liu et al.CVPR 2026
- Read the Room: Video Social Reasoning with Mental-Physical Causal ChainsLixing Niu, Jiapeng Li, Xingping Yu, Xinyi Dong et al.ICLR 2026
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative VideosJiashuo Yu, Yue Wu, Meng Chu, Zhifei Ren et al.ICCV 2025 · 3 citations
- VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric VideosMin Yang, Xinwen Zhang, Jialei Tang, Xin Zhou et al.CVPR 2026
