Zero-Shot Scene Reconstruction from Single Images with Deep Prior Assembly
Junsheng Zhou, Yu-Shen Liu, Zhizhong Han
Abstract
Large language and vision models have been leading a revolution in visual computing. By greatly scaling up sizes of data and model parameters, the large models learn deep priors which lead to remarkable performance in various tasks. In this work, we present deep prior assembly, a novel framework that assembles diverse deep priors from large models for scene reconstruction from single images in a zero-shot manner. We show that this challenging task can be done without extra knowledge but just simply generalizing one deep prior in one sub-task. To this end, we introduce novel methods related to poses, scales, and occlusion parsing which are keys to enable deep priors to work together in a robust way. Deep prior assembly does not require any 3D or 2D data-driven training in the task and demonstrates superior performance in generalizing priors to open-world scenes. We conduct evaluations on various datasets, and report analysis, numerical and visual comparisons with the latest methods to show our superiority. Project page: https://junshengzhou.github.io/DeepPriorAssembly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f815cb1-efee-4b55-a98e-2a1e51b90eedCited by top-tier papers20
- PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion TransformersYuchen Lin, Chenguo Lin, Panwang Pan, Honglei Yan et al.NeurIPS 2025 · 89 citations
- Scenethesis: A Language and Vision Agentic Framework for 3D Scene GenerationLu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding et al.ICLR 2026 · 74 citations
- DiffGS: Functional Gaussian Splatting DiffusionJunsheng Zhou, Weiqi Zhang, Yu-Shen LiuNeurIPS 2024 · 69 citations
- Binocular-Guided 3D Gaussian Splatting with View Consistency for Sparse View SynthesisLiang Han, Junsheng Zhou, Yu-Shen Liu, Zhizhong HanNeurIPS 2024 · 63 citations
- Neural Signed Distance Function Inference through Splatting 3D Gaussians Pulled on Zero-Level SetWenyuan Zhang, Yu-Shen Liu, Zhizhong HanNeurIPS 2024 · 58 citations
Builds on37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Pretrain, Self-train, Distill: A simple recipe for Supersizing 3D ReconstructionKalyan Vasudev Alwala, Abhinav Gupta, Shubham TulsianiCVPR 2022 · 23 citations
- GenPC: Zero-shot Point Cloud Completion via 3D Generative PriorsAn Li, Zhe Zhu, Mingqiang WeiCVPR 2025
- PE3R: Perception-Efficient 3D ReconstructionJie Hu, Shizun Wang, Xinchao WangCVPR 2026 · 9 citations
- Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene ModelingQirui Wu, Denys Iliash, Daniel Ritchie, Manolis Savva et al.ICCV 2025 · 4 citations
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
