PHASE: PHysically-grounded Abstract Social Events for Machine Social Perception
Aviv Netanyahu, Tianmin Shu, Boris Katz, Andrei Barbu, Joshua B. Tenenbaum
Abstract
The ability to perceive and reason about social interactions in the context of physical environments is core to human social intelligence and human-machine cooperation. However, no prior dataset or benchmark has systematically evaluated physically grounded perception of complex social interactions that go beyond short actions, such as high-fiving, or simple group activities, such as gathering. In this work, we create a dataset of physically-grounded abstract social events, PHASE, that resemble a wide range of real-life social interactions by including social concepts such as helping another agent. PHASE consists of 2D animations of pairs of agents moving in a continuous space generated procedurally using a physics engine and a hierarchical planner. Agents have a limited field of view, and can interact with multiple objects, in an environment that has multiple landmarks and obstacles. Using PHASE, we design a social recognition task and a social prediction task. PHASE is validated with human experiments demonstrating that humans perceive rich interactions in the social events, and that the simulated agents behave similarly to humans. As a baseline model, we introduce a Bayesian inverse planning approach, SIMPLE (SIMulation, Planning and Local Estimation), which outperforms state-of-the-art feed-forward neural networks. We hope that PHASE can serve as a difficult new challenge for developing new models that can recognize complex social interactions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of MindYuanfei Wang, Fangwei Zhong, Jing Xu, Yizhou WangICLR 2022 · 103 citations
- AGENT: A Benchmark for Core Psychological ReasoningTianmin Shu, Abhishek Bhandwaldar, Chuang Gan, Kevin A. Smith et al.ICML 2021 · 79 citations
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin et al.AAAI 2025 · 48 citations
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang et al.NeurIPS 2025 · 30 citations
- Interaction Modeling with Multiplex AttentionFan-Yun Sun, Isaac Kauvar, Ruohan Zhang, Jiachen Li et al.NeurIPS 2022 · 29 citations
Builds on3
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory PredictionYingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao et al.ICCV 2019 · 615 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
Related papers
- Modeling dynamic social vision highlights gaps between deep learning and humansKathy Garcia, Emalie McMahon, Colin Conwell, Michael F. Bonner et al.ICLR 2025 · 7 citations
- Virtual Community: An Open World for Humans, Robots, and SocietyQinhong Zhou, Hongxin Zhang, Xiangye Lin, Zheyuan Zhang et al.ICLR 2026 · 12 citations
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan et al.CVPR 2026 · 32 citations
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear et al.ICML 2020 · 88 citations
- Simulation and Retargeting of Complex Multi-Character InteractionsYunbo Zhang, Deepak Gopinath, Yuting Ye, Jessica K. Hodgins et al.SIGGRAPH 2023 · 25 citations
