Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning
YIYAO MA, Kai Chen, Zhongxiang Zhou, Zhuheng Song, Dongsheng Xie, Zelong Tan, Rong Xiong, DOU QI
Abstract
Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object categories remains a significant challenge. In this paper, we present a generalizable deformation learning framework that reconstructs 3D objects by explicitly deforming a category-level shape template to match the target observation. To address complex shape variations between the template and the target, we introduce a geometry-guided feature modeling mechanism. This process first enriches foundation features with template topology to yield a geometry-aware representation, which is then explicitly correlated with the target observation to guide precise deformation. Furthermore, to bridge the disparity between the fixed template and arbitrary target views, we propose a view-adaptive feature aggregation module. This module leverages multi-view template features and their corresponding camera poses to enrich the canonical template representation, ensuring robust feature alignment regardless of the target's perspective. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods in handling large shape variations and diverse viewpoints, exhibiting strong generalization to novel categories and effectively supporting downstream real-world dexterous robotic manipulation tasks. Project homepage: https://GODeform.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd5e6393-fac5-422f-90a1-ce58e5bd147cBuilds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- ICON: Implicit Clothed humans Obtained from NormalsYuliang Xiu, Jinlong Yang, Dimitrios Tzionas, Michael J. BlackCVPR 2022 · 286 citations
- Pixel2Mesh++: Multi-View 3D Mesh Generation via DeformationChao Wen, Yinda Zhang, Zhuwen Li, Yanwei FuICCV 2019 · 279 citations
- Wonder3D: Single Image to 3D Using Cross-Domain DiffusionXiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu et al.CVPR 2024 · 269 citations
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang et al.ICCV 2019 · 218 citations
Related papers
- Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose EstimationXiao Lin, Wenfei Yang, Yuan Gao, Tianzhu ZhangCVPR 2024
- Vision Foundation Model Enables Generalizable Object Pose EstimationKai Chen, Yiyao Ma, Xingyu Lin, Stephen James et al.NeurIPS 2024 · 5 citations
- DTF-Net: Category-Level Pose Estimation and Shape Reconstruction via Deformable Template FieldHaowen Wang, Zhipeng Fan, Zhen Zhao, Zhengping Che et al.ACM MM 2023 · 6 citations
- Autonomous Manipulation Learning for Similar Deformable Objects via Only One DemonstrationYu Ren, Ronghan Chen, Yang CongCVPR 2023
- PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View ReasoningJianqi Chen, Biao Zhang, Xiangjun Tang, Peter WonkaCVPR 2026
