Learning Partonomic 3D Reconstruction from Image Collections
Xiaoqian Ruan, Pei Yu, Dian Jia, Hyeonjeong Park, Peixi Xiong, Wei Tang
Abstract
Reconstructing the 3D shape of an object from a single-view image is a fundamental task in computer vision. Recent advances in differentiable rendering have enabled 3D reconstruction from image collections using only 2D annotations. However, these methods mainly focus on whole-object reconstruction and overlook object partonomy, which is essential for intelligent agents interacting with physical environments. This paper aims at learning partonomic 3D reconstruction from collections of images with only 2D annotations. Our goal is not only to reconstruct the shape of an object from a single-view image but also to decompose the shape into meaningful semantic parts. To handle the expanded solution space and frequent part occlusions in single-view images, we introduce a novel approach that represents, parses, and learns the structural compositionality of 3D objects. This approach comprises: (1) a compact and expressive compositional representation of object geometry, achieved through disentangled modeling of large shape variations, constituent parts, and detailed part deformations as multi-granularity neural fields; (2) a part transformer that recovers precise partonomic geometry and handles occlusions, through effective part-to-pixel grounding and part-to-part relational modeling; and (3) a 2D-supervised learning method that jointly learns the compositional representation and part transformer, by bridging object shape and parts, image synthesis, and differentiable rendering. Extensive experiments on ShapeNetPart, Part-Net, and CUB-200-2011 demonstrate the effectiveness of our approach on both overall and partonomic reconstruction. Code, models, and data are avaliable at https: / / github . com / XiaoqianRuan1 / Partonomic _ Reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec19b419-d8df-4192-b4f8-db8b0c705e31Builds on25
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang et al.ICCV 2019 · 218 citations
- SDF-SRN: Learning Signed Distance 3D Object Reconstruction from Static ImagesChen-Hsuan Lin, Chaoyang Wang, Simon LuceyNeurIPS 2020 · 125 citations
Related papers
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion TransformersYuchen Lin, Chenguo Lin, Panwang Pan, Honglei Yan et al.NeurIPS 2025 · 89 citations
- Discovering 3D Parts from Image CollectionsChun-Han Yao, Wei-Chih Hung, Varun Jampani, Ming-Hsuan YangICCV 2021 · 21 citations
- Unsupervised Volumetric AnimationAliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov, Kyle Olszewski et al.CVPR 2023
- Topologically-Aware Deformation Fields for Single-View 3D ReconstructionShivam Duggal, Deepak PathakCVPR 2022 · 30 citations
