Learning What and Where: Disentangling Location and Identity Tracking Without Supervision
Manuel Traub, Sebastian Otte, Tobias Menge, Matthias Karlbauer, Jannik Thümmel, Martin V. Butz
Abstract
Our brain can almost effortlessly decompose visual data streams into background and salient objects. Moreover, it can anticipate object motion and interactions, which are crucial abilities for conceptual planning and reasoning. Recent object reasoning datasets, such as CATER, have revealed fundamental shortcomings of current vision-based AI systems, particularly when targeting explicit object representations, object permanence, and object reasoning. Here we introduce a self-supervised LOCation and Identity tracking system (Loci), which excels on the CATER tracking challenge. Inspired by the dorsal and ventral pathways in the brain, Loci tackles the binding problem by processing separate, slot-wise encodings of 'what' and 'where'. Loci's predictive coding-like processing encourages active error minimization, such that individual slots tend to encode individual objects. Interactions between objects and object dynamics are processed in the disentangled latent space. Truncated backpropagation through time combined with forward eligibility accumulation significantly speeds up learning and improves memory efficiency. Besides exhibiting superior performance in current benchmarks, Loci effectively extracts objects from video streams and separates them into location and Gestalt components. We believe that this separation offers a representation that will facilitate effective planning and reasoning on conceptual levels. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a41dd02e-617e-48ba-9335-1a877df96e1cCited by top-tier papers13
- Object-Centric Learning for Real-World Videos by Predicting Temporal Feature SimilaritiesAndrii Zadaianchuk, Maximilian Seitzer, Georg MartiusNeurIPS 2023 · 104 citations
- Rotating Features for Object DiscoverySindy Löwe, Phillip Lippe, Francesco Locatello, Max WellingNeurIPS 2023 · 37 citations
- Neural Foundations of Mental Simulation: Future Prediction of Latent Representations on Dynamic ScenesAran Nayebi, Rishi Rajalingham, Mehrdad Jazayeri, Guangyu Robert YangNeurIPS 2023 · 31 citations
- Bridging the Gap to Real-World Object-Centric LearningMaximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow et al.ICLR 2023 · 31 citations
- Entity-Centric Reinforcement Learning for Object Manipulation from PixelsDan Haramati, Tal Daniel, Aviv TamarICLR 2024 · 31 citations
Builds on19
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
Related papers
- CATER: A diagnostic dataset for Compositional Actions & TEmporal ReasoningRohit Girdhar, Deva RamananICLR 2020 · 198 citations
- Look, Remember and Reason: Grounded Reasoning in Videos with Language ModelsApratim Bhattacharyya, Sunny Panchal, Reza Pourreza, Mingu Lee et al.ICLR 2024 · 15 citations
- Does Visual Pretraining Help End-to-End Reasoning?Chen Sun, Calvin Luo, Xingyi Zhou, Anurag Arnab et al.NeurIPS 2023 · 4 citations
- Hopper: Multi-hop Transformer for Spatiotemporal ReasoningHonglu Zhou, Asim Kadav, Farley Lai, Alexandru Niculescu-Mizil et al.ICLR 2021 · 19 citations
- Neural Concept BinderWolfgang Stammer, Antonia Wüst, David Steinmann, Kristian KerstingNeurIPS 2024
