Think-Then-Generate: Structural Chain-of-Thought Reasoning for Consistent 3D Generation
Xinyue Liu, Jin Liu, Hongbo Wang, Ran He, Huaibo Huang
Abstract
Recently, generating 3D assets using visual priors from pretrained diffusion models has shown remarkable results. However, due to the inherent lack of 3D geometric priors in 2D diffusion, the synthesized results often suffer from spatial hallucination and multi-view inconsistency. To address this limitation, we propose Thoughtful3D, a novel framework that enhances 3D content generation quality by introducing structural chain-of-thought (CoT) reasoning to alleviate inconsistent issues and mitigate hallucinations. Specifically, we design a dual-phase structural CoT strategy: (1) 3DBlueprint-CoT explicitly plans the 3D generation process through textual semantic parsing and logical deduction during the initialization phase. (2) 3DRefine-CoT dynamically evaluates latent inconsistencies by analyzing multiple renderings, employing a multi-round iterative refinement mechanism to suppress hallucinations and enhance cross-view consistency. To further promote consistency across views, we propose a Cross-view Semantic Appearance Alignment strategy that enhances multi-view consistency by establishing dynamic geometric associations between the same features from different viewpoints. Extensive experiments demonstrate that Thoughtful3D significantly improves the quality and consistency of generated 3D assets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfbd3e1b-a38a-4aee-8447-deccc0389fc0Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content GenerationHongbo Wang, Jie Cao, Jin Liu, Xiaoqiang Zhou et al.NeurIPS 2024 · 10 citations
- SweetDreamer: Aligning Geometric Priors in 2D diffusion for Consistent Text-to-3DWeiyu Li, Rui Chen, Xuelin Chen, Ping TanICLR 2024 · 155 citations
- Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion PriorsYukang Lin, Haonan Han, Chaoqun Gong, Zunnan Xu et al.ACM MM 2024 · 20 citations
- Retrieval-Augmented Score Distillation for Text-to-3D GenerationJunyoung Seo, Susung Hong, Wooseok Jang, Inès Hyeonsu Kim et al.ICML 2024 · 14 citations
- Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D PriorCheng Chen, Xiaofeng Yang, Fan Yang, Chengzeng Feng et al.CVPR 2024
