Learning About Objects by Learning to Interact with Them
Martin Lohmann, Jordi Salvador, Aniruddha Kembhavi, Roozbeh Mottaghi
Abstract
Much of the remarkable progress in computer vision has been focused around fully supervised learning mechanisms relying on highly curated datasets for a variety of tasks. In contrast, humans often learn about their world with little to no external supervision. Taking inspiration from infants learning from their environment through play and interaction, we present a computational framework to discover objects and learn their physical properties along this paradigm of Learning from Interaction. Our agent, when placed within the near photo-realistic and physics-enabled AI2THOR environment, interacts with its world and learns about objects, their geometric extents and relative masses, without any external guidance. Our experiments reveal that this agent learns efficiently and effectively; not just for objects it has interacted with before, but also for novel instances from seen categories as well as novel object categories. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da6acc5e-5731-40a2-945e-ae77ae6ace5aCited by top-tier papers11
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta et al.ICCV 2021 · 240 citations
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo et al.ICLR 2022 · 119 citations
- Act the Part: Learning Interaction Strategies for Articulated Object Part DiscoverySamir Yitzhak Gadre, Kiana Ehsani, Shuran SongICCV 2021 · 64 citations
- Sound Adversarial Audio-Visual NavigationYinfeng Yu, Wenbing Huang, Fuchun Sun, Changan Chen et al.ICLR 2022 · 49 citations
- Ask4Help: Learning to Leverage an Expert for Embodied TasksKunal Pratap Singh, Luca Weihs, Alvaro Herrasti, Jonghyun Choi et al.NeurIPS 2022 · 33 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Scaling and Benchmarking Self-Supervised Visual Representation LearningPriya Goyal, Dhruv Mahajan, Abhinav Gupta, Ishan MisraICCV 2019 · 429 citations
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall et al.ICCV 2019 · 294 citations
- Visual Reaction: Learning to Play Catch With Your DroneKuo-Hao Zeng, Roozbeh Mottaghi, Luca Weihs, Ali FarhadiCVPR 2020
- Learning Video Object Segmentation From Unlabeled VideosXiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai et al.CVPR 2020
Related papers
- Can Vision Language Models Learn Intuitive Physics from Interaction?Luca M. Schulze Buschoff, Konstantinos Voudouris, Can Demircan, Eric SchulzICML 2026
- Learning Affordance Landscapes for Interaction Exploration in 3D EnvironmentsTushar Nagarajan, Kristen GraumanNeurIPS 2020 · 87 citations
- Unsupervised Discovery of 3D Physical Objects from VideoYilun Du, Kevin A. Smith, Tomer D. Ullman, Joshua B. Tenenbaum et al.ICLR 2021 · 11 citations
- IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor ScenesQi Li, Kaichun Mo, Yanchao Yang, Hang Zhao et al.ICLR 2022 · 9 citations
- Hierarchical Relational InferenceAleksandar Stanic, Sjoerd van Steenkiste, Jürgen SchmidhuberAAAI 2021 · 17 citations
