Robust 3D Shape Reconstruction in Zero-Shot from a Single Image in the Wild
Junhyeong Cho, Kim Youwang, Hunmin Yang, Tae-Hyun Oh
Abstract
Recent monocular 3D shape reconstruction methods have shown promising zero-shot results on object-segmented images without any occlusions. However, their effectiveness is significantly compromised in real-world conditions, due to imperfect object segmentation by off-the-shelf models and the prevalence of occlusions. To effectively address these issues, we propose a unified regression model that integrates segmentation and reconstruction, specifically designed for occlusion-aware 3D shape reconstruction. To facilitate its reconstruction in the wild, we also introduce a scalable data synthesis pipeline that simulates a wide range of variations in objects, occluders, and backgrounds. Training on our synthetic data enables the proposed model to achieve stateof-the-art zero-shot results on real-world images, using significantly fewer parameters than competing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation ModelYukai Shi, Weiyu Li, Zihao Wang, Hongyang Li et al.CVPR 2026 · 17 citations
- Amodal3R: Amodal 3D Reconstruction from Occluded 2D ImagesTianhao Wu, Chuanxia Zheng, Frank Guan, Andrea Vedaldi et al.ICCV 2025 · 9 citations
- Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry RefinementWei Li, Long Ji, Ying Wang, Xiao Wu et al.AAAI 2026
- Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation LearningYIYAO MA, Kai Chen, Zhongxiang Zhou, Zhuheng Song et al.ICML 2026
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Pretrain, Self-train, Distill: A simple recipe for Supersizing 3D ReconstructionKalyan Vasudev Alwala, Abhinav Gupta, Shubham TulsianiCVPR 2022 · 23 citations
- ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic GraspingShun Iwase, Muhammad Zubair Irshad, Katherine Liu, Vitor Guizilini et al.CVPR 2025
- What You Can Reconstruct from a ShadowRuoshi Liu, Sachit Menon, Chengzhi Mao, Dennis Park et al.CVPR 2023
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- ZeroShape: Regression-Based Zero-Shot Shape ReconstructionZixuan Huang, Stefan Stojanov, Anh Thai, Varun Jampani et al.CVPR 2024
