AGENT: A Benchmark for Core Psychological Reasoning
Tianmin Shu, Abhishek Bhandwaldar, Chuang Gan, Kevin A. Smith, Shari Liu, Dan Gutfreund, Elizabeth S. Spelke, Joshua B. Tenenbaum, Tomer D. Ullman
Abstract
For machine agents to successfully interact with humans in real-world settings, they will need to develop an understanding of human mental life. Intuitive psychology, the ability to reason about hidden mental variables that drive observable actions, comes naturally to people: even pre-verbal infants can tell agents from objects, expecting agents to act efficiently to achieve goals given constraints. Despite recent interest in machine agents that reason about other agents, it is not clear if such agents learn or hold the core psychology principles that drive human reasoning. Inspired by cognitive development studies on intuitive psychology, we present a benchmark consisting of a large dataset of procedurally generated 3D animations, AGENT (Action, Goal, Efficiency, coNstraint, uTility), structured around four scenarios (goal preferences, action efficiency, unobserved constraints, and cost-reward trade-offs) that probe key concepts of core intuitive psychology. We validate AGENT with human-ratings, propose an evaluation protocol emphasizing generalization, and compare two strong baselines built on Bayesian inverse planning and a Theory of Mind neural network. Our results suggest that to pass the designed tests of core intuitive psychology at human levels, a model must acquire or have built-in representations of how agents plan, combining utility computations and core knowledge of objects and physics. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68f56d70-e445-490a-b3c0-6c694daff96bCited by top-tier papers16
- ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of MindYuanfei Wang, Fangwei Zhong, Jing Xu, Yizhou WangICLR 2022 · 103 citations
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersKanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. DillonNeurIPS 2021 · 59 citations
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin et al.AAAI 2025 · 48 citations
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang et al.NeurIPS 2025 · 30 citations
- Symmetric Machine Theory of MindMelanie Sclar, Graham Neubig, Yonatan BiskICML 2022 · 22 citations
Builds on3
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersKanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. DillonNeurIPS 2021 · 59 citations
- Few-Shot Bayesian Imitation Learning with Logical Program PoliciesTom Silver, Kelsey R. Allen, Alex K. Lew, Leslie Pack Kaelbling et al.AAAI 2020 · 57 citations
- PHASE: PHysically-grounded Abstract Social Events for Machine Social PerceptionAviv Netanyahu, Tianmin Shu, Boris Katz, Andrei Barbu et al.AAAI 2021 · 44 citations
Related papers
- Neural Reasoning about Agents' Goals, Preferences, and ActionsMatteo Bortoletto, Lei Shi, Andreas BullingAAAI 2024 · 8 citations
- X-VoE: Measuring eXplanatory Violation of Expectation in Physical EventsBo Dai, Linge Wang, Baoxiong Jia, Zeyu Zhang et al.ICCV 2023 · 4 citations
- A Bayesian-Symbolic Approach to Reasoning and Learning in Intuitive PhysicsKai Xu, Akash Srivastava, Dan Gutfreund, Felix Sosa et al.NeurIPS 2021 · 29 citations
- I-PHYRE: Interactive Physical ReasoningShiqian Li, Kewen Wu, Chi Zhang, Yixin ZhuICLR 2024 · 16 citations
- Implicit Intelligence - Evaluating Agents on What Users Don’t SayVed Sirdeshmukh, Marc WetterICML 2026
