Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of others
Kanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. Dillon
Abstract
To achieve human-like common sense about everyday life, machine learning systems must understand and reason about the goals, preferences, and actions of other agents in the environment. By the end of their first year of life, human infants intuitively achieve such common sense, and these cognitive achievements lay the foundation for humans' rich and complex understanding of the mental states of others. Can machines achieve generalizable, commonsense reasoning about other agents like human infants? The Baby Intuitions Benchmark (BIB) 1 challenges machines to predict the plausibility of an agent's behavior based on the underlying causes of its actions. Because BIB's content and paradigm are adopted from developmental cognitive science, BIB allows for direct comparison between human and machine performance. Nevertheless, recently proposed, deep-learning-based agency reasoning models fail to show infant-like reasoning, leaving BIB an open challenge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of MindYuanfei Wang, Fangwei Zhong, Jing Xu, Yizhou WangICLR 2022 · 103 citations
- AGENT: A Benchmark for Core Psychological ReasoningTianmin Shu, Abhishek Bhandwaldar, Chuang Gan, Kevin A. Smith et al.ICML 2021 · 79 citations
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin et al.AAAI 2025 · 48 citations
- Language Models Represent Beliefs of Self and OthersWentao Zhu, Zhining Zhang, Yizhou WangICML 2024 · 24 citations
- Symmetric Machine Theory of MindMelanie Sclar, Graham Neubig, Yonatan BiskICML 2022 · 22 citations
Builds on3
- Decoupling Representation Learning from Reinforcement LearningAdam Stooke, Kimin Lee, Pieter Abbeel, Michael LaskinICML 2021 · 389 citations
- Keep Doing What Worked: Behavior Modelling Priors for Offline Reinforcement LearningNoah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki et al.ICLR 2020 · 299 citations
- AGENT: A Benchmark for Core Psychological ReasoningTianmin Shu, Abhishek Bhandwaldar, Chuang Gan, Kevin A. Smith et al.ICML 2021 · 79 citations
Related papers
- Neural Reasoning about Agents' Goals, Preferences, and ActionsMatteo Bortoletto, Lei Shi, Andreas BullingAAAI 2024 · 8 citations
- BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of MindYuanyuan Mao, Xin Lin, Qin Ni, Liang HeAAAI 2024 · 6 citations
- X-VoE: Measuring eXplanatory Violation of Expectation in Physical EventsBo Dai, Linge Wang, Baoxiong Jia, Zeyu Zhang et al.ICCV 2023 · 4 citations
- VECA: A New Benchmark and Toolkit for General Cognitive DevelopmentKwanyoung Park, Hyunseok Oh, Youngki LeeAAAI 2022
- A Newborn Embodied Turing Test for Comparing Object Segmentation Across Animals and MachinesManju Garimella, Denizhan Pak, Justin N. Wood, Samantha Marie Waters WoodICLR 2024 · 1 citation
