Sherpa3D: Boosting High-Fidelity Text-to-3D Generation via Coarse 3D Prior
Fangfu Liu, Diankun Wu, Yi Wei, Yongming Rao, Yueqi Duan
Abstract
Recently, 3D content creation from text prompts has demonstrated remarkable progress by utilizing 2D and 3D diffusion models. While 3D diffusion models ensure great multi-view consistency, their ability to generate highquality and diverse 3D assets is hindered by the limited 3D data. In contrast, 2D diffusion models find a distillation approach that achieves excellent generalization and rich details without any 3D data. However, 2D lifting methods suffer from inherent view-agnostic ambiguity thereby leading to serious multi-face Janus issues, where text prompts fail to provide sufficient guidance to learn coherent 3D results. Instead of retraining a costly viewpoint-aware model, we study how to fully exploit easily accessible coarse 3D knowledge to enhance the prompts and guide 2D lifting optimization for refinement. In this paper, we propose Sherpa3D, a new text-to-3D framework that achieves high-fidelity, generalizability, and geometric consistency simultaneously. Specifically, we design a pair of guiding strategies derived from the coarse 3D prior generated by the 3D diffusion model: a structural guidance for geometric fidelity and a semantic guidance for 3D coherence. Employing the two types of guidance, the 2D diffusion model enriches the 3D content with diversified and high - quality results. Extensive experiments show the superiority of our Sherpa3D over the state-of-the-art text-to-3D methods in terms of quality and 3D consistency. Project page: https://liuff19.github.io/Sherpa3D/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a394597-ff2a-40b1-9de7-b6805a95b29eCited by top-tier papers17
- Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single ImageKailu Wu, Fangfu Liu, Zhihan Cai, Runjie Yan et al.NeurIPS 2024 · 185 citations
- GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian GenerationChubin Zhang, Hongliang Song, Yi Wei, Chen Yu et al.NeurIPS 2024 · 40 citations
- Dimensionx: Create Any 3D and 4D Scenes From a Single Image With Decoupled Video DiffusionWenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen et al.ICCV 2025 · 11 citations
- One2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single ImagePengfei Wang, Liyi Chen, Zhiyuan Ma, Yanjun Guo et al.ICLR 2026 · 11 citations
- Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content GenerationHongbo Wang, Jie Cao, Jin Liu, Xiaoqiang Zhou et al.NeurIPS 2024 · 10 citations
Builds on46
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- EfficientDreamer: High-Fidelity and Stable 3D Creation via Orthogonal-view Diffusion PriorsZhipeng Hu, Minda Zhao, Chaoyi Zhao, Xinyue Liang et al.CVPR 2024
- Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D PriorCheng Chen, Xiaofeng Yang, Fan Yang, Chengzeng Feng et al.CVPR 2024
- DreamControl: Control-Based Text-to-3D Generation with 3D Self-PriorTianyu Huang, Yihan Zeng, Zhilu Zhang, Wan Xu et al.CVPR 2024
- DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D DataQihao Liu, Yi Zhang, Song Bai, Adam Kortylewski et al.CVPR 2024 · 4 citations
- SweetDreamer: Aligning Geometric Priors in 2D diffusion for Consistent Text-to-3DWeiyu Li, Rui Chen, Xuelin Chen, Ping TanICLR 2024 · 155 citations
