Fully Understanding Generic Objects: Modeling, Segmentation, and Reconstruction
Feng Liu, Luan Tran, Xiaoming Liu
Abstract
Inferring 3D structure of a generic object from a 2D image is a long-standing objective of computer vision. Conventional approaches either learn completely from CADgenerated synthetic data, which have difficulty in inference from real images, or generate 2.5D depth image via intrinsic decomposition, which is limited compared to the full 3D reconstruction. One fundamental challenge lies in how to leverage numerous real 2D images without any 3D ground truth. To address this issue, we take an alternative approach with semi-supervised learning. That is, for a 2D image of a generic object, we decompose it into latent representations of category, shape, albedo, lighting and camera projection matrix, decode the representations to segmented 3D shape and albedo respectively, and fuse these components to render an image well approximating the input image. Using a category-adaptive 3D joint occupancy field (JOF), we show that the complete shape and albedo modeling enables us to leverage real 2D images in both modeling and model fitting. The effectiveness of our approach is demonstrated through superior 3D reconstruction from a single image, being either synthetic or real, and shape segmentation. Code is available at http://cvlab.cse.msu . edu/project-fully3dobject.html.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 666ea574-19d6-46ab-9c0b-601b464d1cdbCited by top-tier papers12
- Learning Clothing and Pose Invariant 3D Shape Representation for Long-Term Person Re-IdentificationFeng Liu, Minchul Kim, ZiAng Gu, Anil Jain et al.ICCV 2023 · 69 citations
- Voxel-based 3D Detection and Reconstruction of Multiple Objects from a Single ImageFeng Liu, Xiaoming LiuNeurIPS 2021 · 43 citations
- Face Relighting with Geometrically Consistent ShadowsAndrew Z. Hou, Michel Sarkis, Ning Bi, Yiying Tong et al.CVPR 2022 · 39 citations
- Multi-View Aggregation Network for Dichotomous Image SegmentationQian Yu, Xiaoqi Zhao, Youwei Pang, Lihe Zhang et al.CVPR 2024 · 13 citations
- S3OD: Towards Generalizable Salient Object Detection with Synthetic DataOrest Kupyn, Hirokatsu Kataoka, Christian RupprechtICLR 2026 · 6 citations
Builds on16
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Texture Fields: Learning Texture Representations in Function SpaceMichael Oechsle, Lars M. Mescheder, Michael Niemeyer, Thilo Strauss et al.ICCV 2019 · 334 citations
- Pixel2Mesh++: Multi-View 3D Mesh Generation via DeformationChao Wen, Yinda Zhang, Zhuwen Li, Yanwei FuICCV 2019 · 279 citations
- BAE-NET: Branched Autoencoder for Shape Co-SegmentationZhiqin Chen, Kangxue Yin, Matthew Fisher, Siddhartha Chaudhuri et al.ICCV 2019 · 153 citations
Related papers
- Pretrain, Self-train, Distill: A simple recipe for Supersizing 3D ReconstructionKalyan Vasudev Alwala, Abhinav Gupta, Shubham TulsianiCVPR 2022 · 23 citations
- De-rendering 3D Objects in the WildFelix Wimbauer, Shangzhe Wu, Christian RupprechtCVPR 2022 · 29 citations
- Shelf-Supervised Mesh Prediction in the WildYufei Ye, Shubham Tulsiani, Abhinav GuptaCVPR 2021
- Unsupervised Learning of Probably Symmetric Deformable 3D Objects From Images in the WildShangzhe Wu, Christian Rupprecht, Andrea VedaldiCVPR 2020
- 3D Shape Reconstruction from 2D Images with Disentangled Attribute FlowXin Wen, Junsheng Zhou, Yu-Shen Liu, Hua Su et al.CVPR 2022 · 47 citations
