Where2Act: From Pixels to Actions for Articulated 3D Objects
Kaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta, Shubham Tulsiani
Abstract
One of the fundamental goals of visual perception is to allow agents to meaningfully interact with their environment. In this paper, we take a step towards that long-term goal – we extract highly localized actionable information related to elementary actions such as pushing or pulling for articulated objects with movable parts. For example, given a drawer, our network predicts that applying a pulling force on the handle opens the drawer. We propose, discuss, and evaluate novel network architectures that given image and depth data, predict the set of actions possible at each pixel, and the regions over articulated parts that are likely to move under the force. We propose a learning-from-interaction framework with an online data sampling strategy that allows us to train the network in simulation (SAPIEN) and generalizes across categories. Check the website for code and data release.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23aa9ee3-2779-436e-bd7d-3d4a7273da70Cited by top-tier papers74
- ATISS: Autoregressive Transformers for Indoor Scene SynthesisDespoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis et al.NeurIPS 2021 · 293 citations
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and ManipulationJiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An et al.NeurIPS 2024 · 154 citations
- A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape RepresentationJiteng Mu, Weichao Qiu, Adam Kortylewski, Alan L. Yuille et al.ICCV 2021 · 138 citations
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo et al.ICLR 2022 · 119 citations
- PARIS: Part-level Reconstruction and Motion Analysis for Articulated ObjectsJiayi Liu, Ali Mahdavi-Amiri, Manolis SavvaICCV 2023 · 103 citations
Builds on7
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 194 citations
- Learning Affordance Landscapes for Interaction Exploration in 3D EnvironmentsTushar Nagarajan, Kristen GraumanNeurIPS 2020 · 87 citations
- Learning About Objects by Learning to Interact with ThemMartin Lohmann, Jordi Salvador, Aniruddha Kembhavi, Roozbeh MottaghiNeurIPS 2020 · 19 citations
- SAPIEN: A SimulAted Part-Based Interactive ENvironmentFanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia et al.CVPR 2020
- Category-Level Articulated Object Pose EstimationXiaolong Li, He Wang, Li Yi, Leonidas J. Guibas et al.CVPR 2020
Related papers
- Pushing It Out of the Way: Interactive Visual NavigationKuo-Hao Zeng, Luca Weihs, Ali Farhadi, Roozbeh MottaghiCVPR 2021
- Understanding Object Dynamics for Interactive Image-to-Video SynthesisAndreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn OmmerCVPR 2021
- Tactile Sketch SaliencyJianbo Jiao, Ying Cao, Manfred Lau, Rynson W. H. LauACM MM 2020 · 4 citations
- Deep Compliant ControlSeunghwan Lee, Phil Sik Chang, Jehee LeeSIGGRAPH 2022 · 15 citations
- Learning Intuitive Policies Using Action FeaturesMingwei Ma, Jizhou Liu, Samuel Sokota, Max Kleiman-Weiner et al.ICML 2023 · 4 citations
