A Newborn Embodied Turing Test for Comparing Object Segmentation Across Animals and Machines
Manju Garimella, Denizhan Pak, Justin N. Wood, Samantha Marie Waters Wood
Abstract
Newborn brains rapidly learn to solve challenging object perception tasks, including segmenting objects from backgrounds and recognizing objects across new viewing situations. Conversely, modern machine learning (ML) algorithms are "data hungry," requiring more training data than brains to reach similar performance levels. How do we close this learning gap between brains and machines? Here, we introduce a new benchmark-a Newborn Embodied Turing Test (NETT) for object segmentation-in which newborn animals and machines are raised in the same environments and tested with the same tasks, permitting direct comparison of their learning. First, newborn chicks were raised in controlled environments containing a single object rotating on a single background, then their recognition performance was tested across new backgrounds and viewpoints. Second, we performed "digital twin" experiments in which artificial agents were reared and tested in virtual environments that mimicked the rearing and testing conditions of the chicks. We inserted a variety of ML "brains" into the artificial agents and measured whether those algorithms learned common object recognition behavior as chicks. All newborn chicks solved this one-shot object segmentation task, successfully learning background-invariant object representations that generalized across new backgrounds and viewpoints. In contrast, none of the artificial agents solved the task, instead learning background-dependent representations that failed to generalize across new backgrounds and viewpoints. This digital twin design exposes core limitations in current ML algorithms in developing brain-like object perception. Our NETT is publicly available for comparing ML algorithms with newborn chicks. We argue that NETT benchmarks can help researchers build embodied AI systems that learn as efficiently and robustly as newborn brains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- The Unsurprising Effectiveness of Pre-Trained Vision Models for ControlSimone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, Abhinav GuptaICML 2022 · 233 citations
- Self-supervised learning through the eyes of a childA. Emin Orhan, Vaibhav V. Gupta, Brenden M. LakeNeurIPS 2020 · 119 citations
Related papers
- Are Vision Transformers More Data Hungry Than Newborn Visual Systems?Lalit Pandey, Samantha M. W. Wood, Justin N. WoodNeurIPS 2023 · 24 citations
- Online Reasoning Video Segmentation with Just-in-Time Digital TwinsYiqing Shen, Bohan Liu, Chenjia Li, Lalithkumar Seenivasan et al.ICCV 2025 · 7 citations
- The 3D-PC: a benchmark for visual perspective taking in humans and machinesDrew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj et al.ICLR 2025
- Navigation Turing Test (NTT): Learning to Evaluate Human-Like NavigationSam Devlin, Raluca Georgescu, Ida Momennejad, Jaroslaw Rzepecki et al.ICML 2021 · 27 citations
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersKanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. DillonNeurIPS 2021 · 59 citations
