DFGAP: Towards Depth-Free Cross-Category GAParts Perception via Uncertainty-Quantified Modeling
Xueyu Yuan, Jiarui Zhang, Jiangqi Song, Liu Liu, Li Zhang, Dan Guo, Richang Hong, Meng Wang
Abstract
Cross-category object perception is one of the essential upstream tasks for generelizable robot object interaction and manipulation. Recently, an increasing number of researchers are focusing on investigating visual Generalizable and Actionable Parts understanding at cross-category level perception. However, these works are built upon the RGB-D or point cloud input, that relies on the depth information capture. Under the circumstances of limited depth camera performance, e.g. transparent or light absorbing material, perception algorithms that do not require depth information are urgently needed. In this paper, we propose DFGAP, a novel depth-free framework for RGB-based GAParts segmentation and pose estimation. Specifically, we independently model the ill-pose problems from the absence of depth for GAPart segmentation and pose estimation, by clearly quantifying the pixel-wise segmentation probability and relative depth. We reduce the uncertainty and benefit learning in these two tasks. The experimental results demonstrate the superior performance and robustness of our DFGAP. Our work provides a new research paradigm in GAParts perception. We believe that our work has the enormous potential to be applied in many areas of embodied AI system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6185be53-a8f7-4a4e-aef9-dd9223a70f4bBuilds on20
- SoftGroup for 3D Instance Segmentation on Point CloudsThang Vu, Kookhoi Kim, Tung Minh Luu, Thanh Xuan Nguyen et al.CVPR 2022 · 251 citations
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta et al.ICCV 2021 · 240 citations
- GPV-Pose: Category-level Object Pose Estimation via Geometry-guided Point-wise VotingYan Di, Ruida Zhang, Zhiqiang Lou, Fabian Manhardt et al.CVPR 2022 · 141 citations
- Rolling-Unet: Revitalizing MLP's Ability to Efficiently Extract Long-Distance Dependencies for Medical Image SegmentationYutong Liu, Haijiang Zhu, Mengting Liu, Huaiyuan Yu et al.AAAI 2024 · 136 citations
- Generative Category-level Object Pose Estimation via Diffusion ModelsJiyao Zhang, Mingdong Wu, Hao DongNeurIPS 2023 · 65 citations
Related papers
- Generalizable and Actionable Parts Pose Estimation with Symmetry Annotation-Free Learning Strategywenxiao chen, Xueyu Yuan, Liu Liu, Di Wu et al.ICML 2026
- GAPartNet: Cross-Category Domain-Generalizable Object Perception and Manipulation via Generalizable and Actionable PartsHaoran Geng, Helin Xu, Chengyang Zhao, Chao Xu et al.CVPR 2023
- KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose TrackingLiu Liu, Anran Huang, Qi Wu, Dan Guo et al.AAAI 2024 · 7 citations
- GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D ReconstructionLi Zhang, Mingliang Xu, Jianan Wang, Qiaojun Yu et al.CVPR 2025
- Adaptive Articulated Object Manipulation on the Fly with Foundation Model Reasoning and Part GroundingXiaojie Zhang, Yuanfei Wang, Ruihai Wu, Kunqi Xu et al.ICCV 2025 · 2 citations
