AGENT: A Benchmark for Core Psychological Reasoning
Tianmin Shu, Abhishek Bhandwaldar, Chuang Gan, Kevin A. Smith, Shari Liu, Dan Gutfreund, Elizabeth S. Spelke, Joshua B. Tenenbaum, Tomer D. Ullman
摘要
For machine agents to successfully interact with humans in real-world settings, they will need to develop an understanding of human mental life. Intuitive psychology, the ability to reason about hidden mental variables that drive observable actions, comes naturally to people: even pre-verbal infants can tell agents from objects, expecting agents to act efficiently to achieve goals given constraints. Despite recent interest in machine agents that reason about other agents, it is not clear if such agents learn or hold the core psychology principles that drive human reasoning. Inspired by cognitive development studies on intuitive psychology, we present a benchmark consisting of a large dataset of procedurally generated 3D animations, AGENT (Action, Goal, Efficiency, coNstraint, uTility), structured around four scenarios (goal preferences, action efficiency, unobserved constraints, and cost-reward trade-offs) that probe key concepts of core intuitive psychology. We validate AGENT with human-ratings, propose an evaluation protocol emphasizing generalization, and compare two strong baselines built on Bayesian inverse planning and a Theory of Mind neural network. Our results suggest that to pass the designed tests of core intuitive psychology at human levels, a model must acquire or have built-in representations of how agents plan, combining utility computations and core knowledge of objects and physics. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of MindYuanfei Wang, Fangwei Zhong, Jing Xu, Yizhou WangICLR 2022 · 被引用 103 次
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersKanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. DillonNeurIPS 2021 · 被引用 59 次
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin 等AAAI 2025 · 被引用 48 次
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang 等NeurIPS 2025 · 被引用 30 次
- Symmetric Machine Theory of MindMelanie Sclar, Graham Neubig, Yonatan BiskICML 2022 · 被引用 22 次
它引用的顶会 Paper3
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersKanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. DillonNeurIPS 2021 · 被引用 59 次
- Few-Shot Bayesian Imitation Learning with Logical Program PoliciesTom Silver, Kelsey R. Allen, Alex K. Lew, Leslie Pack Kaelbling 等AAAI 2020 · 被引用 57 次
- PHASE: PHysically-grounded Abstract Social Events for Machine Social PerceptionAviv Netanyahu, Tianmin Shu, Boris Katz, Andrei Barbu 等AAAI 2021 · 被引用 44 次
相关 Paper
- Neural Reasoning about Agents' Goals, Preferences, and ActionsMatteo Bortoletto, Lei Shi, Andreas BullingAAAI 2024 · 被引用 8 次
- X-VoE: Measuring eXplanatory Violation of Expectation in Physical EventsBo Dai, Linge Wang, Baoxiong Jia, Zeyu Zhang 等ICCV 2023 · 被引用 4 次
- A Bayesian-Symbolic Approach to Reasoning and Learning in Intuitive PhysicsKai Xu, Akash Srivastava, Dan Gutfreund, Felix Sosa 等NeurIPS 2021 · 被引用 29 次
- I-PHYRE: Interactive Physical ReasoningShiqian Li, Kewen Wu, Chi Zhang, Yixin ZhuICLR 2024 · 被引用 16 次
- Implicit Intelligence - Evaluating Agents on What Users Don’t SayVed Sirdeshmukh, Marc WetterICML 2026
