GMAP: Generalized Manipulation of Articulated Objects in Robotic Using Pre-trained Model
Hongliang Zeng, Ping Zhang, Fang Li, Qinpeng Yi, Tingyu Ye, Jiahua Wang
Abstract
Perception and interaction with articulated objects present a unique challenge for service robots. Although recent research has emphasized understanding articulated shapes and affordance proposals, existing methods only address isolated aspects, failing to develop comprehensive strategies for robotic perception and manipulation of articulated objects. To bridge this gap, we propose GMAP, which systematically integrates the entire process from command to perception and manipulation. Specifically, we first perform precise part-level segmentation of the object and identify the geometric and kinematic parameters of articulated joints. Then, by evaluating point-level affordance proposals, we determine the interaction poses for the robot's end-effector. Finally, the robot's execution trajectory is dynamically computed by combining commands with joint parameters and interaction points. Additionally, a key innovation of GMAP is addressing the scarcity of annotated data. We designed a multi-scale point cloud feature extraction module and introduced pre-training and fine-tuning techniques, significantly enhancing the generalization capability of the perception model. Extensive experiments demonstrate that GMAP achieves state-of-the-art (SOTA) performance in both the perception and manipulation of articulated objects and adapts to real-world scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44cda333-3b8d-4b7d-8cf7-32030b597c65Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Self-Supervised Pretraining of 3D Features on any Point-CloudZaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan MisraICCV 2021 · 333 citations
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta et al.ICCV 2021 · 240 citations
Related papers
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo et al.ICLR 2022 · 119 citations
- GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance GuidanceShuaihang Yuan, Hao Huang, Yu Hao, Congcong Wen et al.NeurIPS 2024 · 42 citations
- Command-driven Articulated Object Understanding and ManipulationRuihang Chu, Zhengzhe Liu, Xiaoqing Ye, Xiao Tan et al.CVPR 2023
- Adaptive Articulated Object Manipulation on the Fly with Foundation Model Reasoning and Part GroundingXiaojie Zhang, Yuanfei Wang, Ruihai Wu, Kunqi Xu et al.ICCV 2025 · 2 citations
- LASO: Language-Guided Affordance Segmentation on 3D ObjectYicong Li, Na Zhao, Junbin Xiao, Chun Feng et al.CVPR 2024
