Unsupervised Object Keypoint Learning using Local Spatial Predictability
Anand Gopalakrishnan, Sjoerd van Steenkiste, Jürgen Schmidhuber
Abstract
We propose PermaKey, a novel approach to representation learning based on object keypoints. It leverages the predictability of local image regions from spatial neighborhoods to identify salient regions that correspond to object parts, which are then converted to keypoints. Unlike prior approaches, it utilizes predictability to discover object keypoints, an intrinsic property of objects. This ensures that it does not overly bias keypoints to focus on characteristics that are not unique to objects, such as movement, shape, colour etc. We demonstrate the efficacy of PermaKey on Atari where it learns keypoints corresponding to the most salient object parts and is robust to certain visual distractors. Further, on downstream RL tasks in the Atari domain we demonstrate how agents equipped with our keypoints outperform those using competing alternatives, even on challenging environments with moving backgrounds or distractor objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3520cc1-6a28-4a68-9cea-1e71816e207eCited by top-tier papers8
- VAEL: Bridging Variational Autoencoders and Probabilistic Logic ProgrammingEleonora Misino, Giuseppe Marra, Emanuele SansoneNeurIPS 2022 · 38 citations
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 35 citations
- Self-Supervised Attention-Aware Reinforcement LearningHaiping Wu, Khimya Khetarpal, Doina PrecupAAAI 2021 · 33 citations
- 3D Implicit Transporter for Temporally Consistent Keypoint DiscoveryChengliang Zhong, Yuhang Zheng, Yupeng Zheng, Hao Zhao et al.ICCV 2023 · 23 citations
- Contrastive Training of Complex-Valued Autoencoders for Object DiscoveryAleksandar Stanic, Anand Gopalakrishnan, Kazuki Irie, Jürgen SchmidhuberNeurIPS 2023 · 21 citations
Builds on6
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 322 citations
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun et al.ICLR 2020 · 276 citations
- Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over ModulesSarthak Mittal, Alex Lamb, Anirudh Goyal, Vikram Voleti et al.ICML 2020 · 73 citations
Related papers
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- Entropy-driven Unsupervised Keypoint Representation Learning in VideosAli Younes, Simone Schaub-Meyer, Georgia ChalvatzakiICML 2023 · 1 citation
- Learning Task-relevant Representations for Generalization via Characteristic Functions of Reward Sequence DistributionsRui Yang, Jie Wang, Zijie Geng, Mingxuan Ye et al.KDD 2022 · 13 citations
- Learning Robust Representations with Long-Term Information for Generalization in Visual Reinforcement LearningRui Yang, Jie Wang, Qijie Peng, Ruibo Guo et al.ICLR 2025
- Unsupervised Visual Attention and Invariance for Reinforcement LearningXudong Wang, Long Lian, Stella X. YuCVPR 2021
