REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment
Haonan Han, Rui Yang, Huan Liao, Jiankai Xing, Zunnan Xu, Xiaoming Yu, Junwei Zha, Xiu Li, Wanhua Li
Abstract
Traditional image-to-3D models often struggle with scenes containing multiple objects due to biases and occlusion complexities. To address this challenge, we present REPARO, a novel approach for compositional 3D asset generation from single images. REPARO employs a two-step process: first, it extracts individual objects from the scene and reconstructs their 3D meshes using off-the-shelf image-to-3D models; then, it optimizes the layout of these meshes through differentiable rendering techniques, ensuring coherent scene composition. By integrating optimal transport-based long-range appearance loss term and high-level semantic loss term in the differentiable rendering, REPARO can effectively recover the layout of 3D assets. The proposed method can significantly enhance object independence, detail accuracy, and overall scene coherence. Extensive evaluation of multi-object scenes demonstrates that our REPARO offers a comprehensive approach to address the complexities of multi-object 3D scene generation from single images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 950907ab-e0e2-4f42-a76e-a304cc3f3497Cited by top-tier papers20
- PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion TransformersYuchen Lin, Chenguo Lin, Panwang Pan, Honglei Yan et al.NeurIPS 2025 · 89 citations
- LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPSWanhua Li, Yujie Zhao, Minghan Qin, Yang Liu et al.NeurIPS 2025 · 54 citations
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial ReasoningJinkun Hao, Naifu Liang, Zhen Luo, Xudong Xu et al.NeurIPS 2025 · 21 citations
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation ModelYukai Shi, Weiyu Li, Zihao Wang, Hongyang Li et al.CVPR 2026 · 17 citations
- PAT3D: Physics-Augmented Text-to-3D Scene GenerationGuying Lin, Kemeng Huang, Michael Liu, Ruihan Gao et al.ICLR 2026 · 14 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- MIDI: Multi-Instance Diffusion for Single Image to 3D Scene GenerationZehuan Huang, Yuan-Chen Guo, Xingqiao An, Yunhan Yang et al.CVPR 2025
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageZe-Xin Yin, Liu Liu, Xinjie wang, Wei Sui et al.CVPR 2026 · 10 citations
- DepR: Depth Guided Single-View Scene Reconstruction with Instance-Level DiffusionQingcheng Zhao, Xiang Zhang, Haiyang Xu, Zeyuan Chen et al.ICCV 2025 · 3 citations
- Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion ModelsZiyi Wu, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson et al.NeurIPS 2024 · 32 citations
- Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion PriorsYukang Lin, Haonan Han, Chaoqun Gong, Zunnan Xu et al.ACM MM 2024 · 20 citations
