3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
Ze-Xin Yin, Liu Liu, Xinjie wang, Wei Sui, Zhizhong Su, Jian Yang, Jin Xie
摘要
We introduce 3D-Fixer, a novel generalizable and efficient scheme for single-image to compositional 3D scene generation. Unlike existing feed-forward frameworks that lack generalization ability in open-set scenarios due to the limited dataset, or divide-and-conquer frameworks that suffer from slow inference or accumulated registration errors during layout alignment, 3D-Fixer extends pre-trained object-level 3D generation priors to perform in-place completion on the single-view estimated geometry, eliminating the need for pose alignment while preserving feed-forward efficiency. At its core, 3D-Fixer introduces a coarse-to-fine scheme to accurately determine the completion boundary and generate high quality completion 3D asset based on the single-view estimated fragmented geometry. Also, we design a dual-branch conditioning network that integrates 2D and 3D contextual information to guide the pre-trained object generation priors for in-place completion. Furthermore, we introduce the Occlusion-Robust Feature Alignment strategy, which employs feature distillation to stabilize the training of the generative priors under occlusion scenarios. Existing scene-level dataset, either suffering from limited scale or lacking accurate per-instance ground truth, severely restricting the development of scene generation approaches. Therefore, we constructed the large-scale scene-level dataset, featuring over 110K diverse scenes and 3M images with complete 3D asset ground truth and accurate placement annotation. Experiments demonstrate that 3D-Fixer achieves state-of-the-art geometric accuracy while maintaining an inference speed comparable to feed-forward estimation methods, vastly outperforming iterative optimization approaches. Our dataset and trained models will be publicly available upon acceptance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation ModelYukai Shi, Weiyu Li, Zihao Wang, Hongyang Li 等CVPR 2026 · 被引用 17 次
- Pano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic ImageZidian Qiu, Ancong WuCVPR 2026 · 被引用 1 次
- REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout AlignmentHaonan Han, Rui Yang, Huan Liao, Jiankai Xing 等ICCV 2025 · 被引用 4 次
- CC3D: Layout-Conditioned Generation of Compositional 3D ScenesSherwin Bahmani, Jeong Joon Park, Despoina Paschalidou, Xingguang Yan 等ICCV 2023 · 被引用 66 次
- WorldGrow: Generating Infinite 3D WorldSikuang Li, Chen Yang, Jiemin Fang, Taoran Yi 等AAAI 2026 · 被引用 10 次
