Unsupervised Discovery of 3D Physical Objects from Video
Yilun Du, Kevin A. Smith, Tomer D. Ullman, Joshua B. Tenenbaum, Jiajun Wu
Abstract
We study the problem of unsupervised physical object discovery. While existing frameworks aim to decompose scenes into 2D segments based off each object's appearance, we explore how physics, especially object interactions, facilitates disentangling of 3D geometry and position of objects from video, in an unsupervised manner. Drawing inspiration from developmental psychology, our Physical Object Discovery Network (POD-Net) uses both multi-scale pixel cues and physical motion cues to accurately segment observable and partially occluded objects of varying sizes, and infer properties of those objects. Our model reliably segments objects on both synthetic and real scenes. The discovered object properties can also be used to reason about physical events. INTRODUCTION From early in development, infants impose structure on their world. When they look at a scene, infants do not perceive simply an array of colors. Instead, they scan the scene and organize the world into objects that obey certain physical expectations, like traveling along smooth paths or not winking in and out of existence (Spelke & Kinzler, 2007; Spelke et al., 1992) . Here we take two ideas from human, and particularly infant, perception for helping artificial agents learn about object properties: that coherent object motion constrains expectations about future object states, and that foveation patterns allow people to scan both small or far-away and large or close-up objects in the same scene. Motion is particularly crucial in the early ability to segment a scene into individual objects. For instance, infants perceive two patches moving together as a single object, even though they look perceptually distinct to adults (Kellman & Spelke, 1983) . This segmentation from motion even leads young children to expect that if a toy resting on a block is picked up, both the block and the toy will move up as if they are a single object. This suggests that artificial systems that learn to segment the world could be usefully constrained by the principle that there are objects that move in regular ways. In addition, human vision exhibits foveation patterns, where only a local patch of a scene is often visible at once. This allows people to focus on objects that are otherwise small on the retina, but also stitch together different glimpses of larger objects into a coherent whole. We propose the Physical Object Discovery Network (POD-Net), a self-supervised model that learns to extract object-based scene representations from videos using motion cues. POD-Net links a visual generative model with a dynamics model in which objects persist and move smoothly. The visual generative model factors an object-based scene decompositions across local patches, then aggregates those local patches into a global segmentation. The link between the visual model and the dynamics model constrains the discovered representations to be usable to predict future world states. POD-Net thus produces more stable image segmentations than other self-supervised segmentation models, especially in challenging conditions such as when objects occlude each other (Figure 1 ). We test how well POD-Net performs image segmentation and object discovery on two datasets: one made from ShapeNet objects (Chang et al., 2015) , and one from real-world images. We find that POD-Net outperforms recent self-supervised image segmentation models that use regular foregroundbackground relationships (Greff et al., 2019) or assume that images are composable into object-like parts (Burgess et al., 2019). Finally, we show that the representations learned by POD-Net can be used to support reasoning in a task that requires identifying scenes with physically implausible events
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ec835f1-b64f-43ac-aff1-7b64f73d1b42Cited by top-tier papers13
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone et al.ICLR 2022 · 290 citations
- Simple Unsupervised Object-Centric Learning for Complex and Naturalistic VideosGautam Singh, Yi-Fu Wu, Sungjin AhnNeurIPS 2022 · 182 citations
- Object-Centric Slot DiffusionJindong Jiang, Fei Deng, Gautam Singh, Sungjin AhnNeurIPS 2023 · 106 citations
- Attention over Learned Object Embeddings Enables Complex Visual ReasoningDavid Ding, Felix Hill, Adam Santoro, Malcolm Reynolds et al.NeurIPS 2021 · 87 citations
- Rotating Features for Object DiscoverySindy Löwe, Phillip Lippe, Francesco Locatello, Max WellingNeurIPS 2023 · 37 citations
Builds on4
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying et al.ICML 2020 · 1,439 citations
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall et al.ICCV 2019 · 294 citations
- Anchor Diffusion for Unsupervised Video Object SegmentationZhao Yang, Qiang Wang, Luca Bertinetto, Song Bai et al.ICCV 2019 · 127 citations
- DOPS: Learning to Detect 3D Objects and Predict Their 3D ShapesMahyar Najibi, Guangda Lai, Abhijit Kundu, Zhichao Lu et al.CVPR 2020
Related papers
- Learning Physical Graph Representations from Visual ScenesDaniel Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li et al.NeurIPS 2020 · 88 citations
- Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual GroupingLong Lian, Zhirong Wu, Stella X. YuCVPR 2023
- Hierarchical Relational InferenceAleksandar Stanic, Sjoerd van Steenkiste, Jürgen SchmidhuberAAAI 2021 · 17 citations
- Object Concepts Emerge from MotionHaoqian Liang, Xiaohui Wang, Zhichao Li, Ya Yang et al.NeurIPS 2025
- PARTS: Unsupervised segmentation with slots, attention and independence maximizationDaniel Zoran, Rishabh Kabra, Alexander Lerchner, Danilo J. RezendeICCV 2021 · 53 citations
