Exploiting Spatial Invariance for Scalable Unsupervised Object Tracking
Eric Crawford, Joelle Pineau
Abstract
The ability to detect and track objects in the visual world is a crucial skill for any intelligent agent, as it is a necessary precursor to any object-level reasoning process. Moreover, it is important that agents learn to track objects without supervision (i.e. without access to annotated training videos) since this will allow agents to begin operating in new environments with minimal human assistance. The task of learning to discover and track objects in videos, which we call unsupervised object tracking, has grown in prominence in recent years; however, most architectures that address it still struggle to deal with large scenes containing many objects. In the current work, we propose an architecture that scales well to the large-scene, many-object setting by employing spatially invariant computations (convolutions and spatial attention) and representations (a spatially local object specification scheme). In a series of experiments, we demonstrate a number of attractive features of our architecture; most notably, that it outperforms competing methods at tracking objects in cluttered scenes with many objects, and that it can generalize well to videos that are larger and/or contain more objects than videos encountered during training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3708b21b-d849-4d1a-af1a-b39de0ac1f18Cited by top-tier papers27
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone et al.ICLR 2022 · 290 citations
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun et al.ICLR 2020 · 276 citations
- Simple Unsupervised Object-Centric Learning for Complex and Naturalistic VideosGautam Singh, Yi-Fu Wu, Sungjin AhnNeurIPS 2022 · 182 citations
- Illiterate DALL-E Learns to ComposeGautam Singh, Fei Deng, Sungjin AhnICLR 2022 · 182 citations
- SCALOR: Generative World Models with Scalable Object RepresentationsJindong Jiang, Sepehr Janghorbani, Gerard de Melo, Sungjin AhnICLR 2020 · 152 citations
Related papers
- Object-Centric Multiple Object TrackingZixu Zhao, Jiaze Wang, Max Horn, Yizhuo Ding et al.ICCV 2023 · 10 citations
- VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAEHaonan Yu, Wei XuICLR 2024 · 1 citation
- VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object TrackingZekun Qian, Ruize Han, Junhui Hou, Linqi Song et al.ICCV 2025 · 3 citations
- Self-Supervised Object Detection from Egocentric VideosPeri Akiva, Jing Huang, Kevin J. Liang, Rama Kovvuri et al.ICCV 2023 · 9 citations
- Unsupervised Multi-Object Segmentation by Predicting Probable Motion PatternsLaurynas Karazija, Subhabrata Choudhury, Iro Laina, Christian Rupprecht et al.NeurIPS 2022 · 24 citations
