A Newborn Embodied Turing Test for Comparing Object Segmentation Across Animals and Machines
Manju Garimella, Denizhan Pak, Justin N. Wood, Samantha Marie Waters Wood
摘要
Newborn brains rapidly learn to solve challenging object perception tasks, including segmenting objects from backgrounds and recognizing objects across new viewing situations. Conversely, modern machine learning (ML) algorithms are "data hungry," requiring more training data than brains to reach similar performance levels. How do we close this learning gap between brains and machines? Here, we introduce a new benchmark-a Newborn Embodied Turing Test (NETT) for object segmentation-in which newborn animals and machines are raised in the same environments and tested with the same tasks, permitting direct comparison of their learning. First, newborn chicks were raised in controlled environments containing a single object rotating on a single background, then their recognition performance was tested across new backgrounds and viewpoints. Second, we performed "digital twin" experiments in which artificial agents were reared and tested in virtual environments that mimicked the rearing and testing conditions of the chicks. We inserted a variety of ML "brains" into the artificial agents and measured whether those algorithms learned common object recognition behavior as chicks. All newborn chicks solved this one-shot object segmentation task, successfully learning background-invariant object representations that generalized across new backgrounds and viewpoints. In contrast, none of the artificial agents solved the task, instead learning background-dependent representations that failed to generalize across new backgrounds and viewpoints. This digital twin design exposes core limitations in current ML algorithms in developing brain-like object perception. Our NETT is publicly available for comparing ML algorithms with newborn chicks. We argue that NETT benchmarks can help researchers build embodied AI systems that learn as efficiently and robustly as newborn brains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- The Unsurprising Effectiveness of Pre-Trained Vision Models for ControlSimone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, Abhinav GuptaICML 2022 · 被引用 233 次
- Self-supervised learning through the eyes of a childA. Emin Orhan, Vaibhav V. Gupta, Brenden M. LakeNeurIPS 2020 · 被引用 119 次
相关 Paper
- Are Vision Transformers More Data Hungry Than Newborn Visual Systems?Lalit Pandey, Samantha M. W. Wood, Justin N. WoodNeurIPS 2023 · 被引用 24 次
- Online Reasoning Video Segmentation with Just-in-Time Digital TwinsYiqing Shen, Bohan Liu, Chenjia Li, Lalithkumar Seenivasan 等ICCV 2025 · 被引用 7 次
- The 3D-PC: a benchmark for visual perspective taking in humans and machinesDrew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj 等ICLR 2025
- Navigation Turing Test (NTT): Learning to Evaluate Human-Like NavigationSam Devlin, Raluca Georgescu, Ida Momennejad, Jaroslaw Rzepecki 等ICML 2021 · 被引用 27 次
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersKanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. DillonNeurIPS 2021 · 被引用 59 次
