Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image
Kailu Wu, Fangfu Liu, Zhihan Cai, Runjie Yan, Hanyang Wang, Yating Hu, Yueqi Duan, Kaisheng Ma
Abstract
In this work, we introduce Unique3D, a novel image-to-3D framework for efficiently generating high-quality 3D meshes from single-view images, featuring state-of-the-art generation fidelity and strong generalizability. Previous methods based on Score Distillation Sampling (SDS) can produce diversified 3D results by distilling 3D knowledge from large 2D diffusion model, but they usually suffer from long per-case optimization time with inconsistent issues. Recent works address the problem and generate better 3D results either by finetuning a multi-view diffusion model or training a fast feed-forward model. However, they still lack intricate textures and complex geometries due to inconsistency and limited gener-38th Conference on Neural Information Processing Systems (NeurIPS 2024).
ated resolution. To simultaneously achieve high fidelity, consistency, and efficiency in single image-to-3D, we propose a novel framework Unique3D that includes a multi-view diffusion model with a corresponding normal diffusion model to generate multi-view images with their normal maps, a multi-level upscale process to progressively improve the resolution of generated orthographic multi-views, as well as an instant and consistent mesh reconstruction algorithm called ISOMER, which fully integrates the color and geometric priors into mesh results. Extensive experiments demonstrate that our Unique3D significantly outperforms other image-to-3D baselines in terms of geometric and textural details. Project page: https://wukailu.github.io/Unique3D/.
Recently, the rapid development of diffusion models [12,52,39] has opened up new perspectives for 3D content creation. Powered by the strong prior of 2D image diffusion models, DreamFusion [36] proposes Score Distillation Sampling (SDS) to address the limitation of 3D data by distilling 3D knowledge from 2D diffusions [41], inspiring the progress of SDS-based 2D lifting methods [20,37,59,24,5]. Despite their diversified compelling results, they usually suffer from long per-case optimization time for hours, poor geometry, and inconsistent issues (e.g., , Janus problem [36]), thus not practical for real-world applications. To overcome the problems, a series of works leverage largerscale open-world 3D datasets [7,4,6] either to fine-tune a multi-view diffusion model [29,28,57] and recover the 3D shapes from the generated multi-view images or train a large reconstruction model (LRM) [13,64,60,63] by directly mapping image tokens into 3D representations (e.g., , triplane or 3D Gaussian [16]). However, due to local inconsistency in mesh optimization [29,65] and limited resolution of the generative process with expensive computational overhead [13,63], they struggle to produce intricate textures and complex geometric details with high resolution.
In this paper, we present a novel image-to-3D framework for efficient 3D mesh generation, coined Unique3D, to address the above challenges and simultaneously achieve high-fidelity, consistency, and generalizability. Given an input image, Unique3D first generates orthographic multi-view images from a multi-view diffusion model. Then we introduce a multi-level upscale strategy to progressively improve the resolution of generated multi-view images with their corresponding normal maps from a normal diffusion model. Finally, we propose an instant and consistent mesh reconstruction (ISOMER) algorithm to reconstruct high-quality 3D meshes from the multiple RGB images and normal maps, which fully integrates the color and geometric priors into mesh results. Both diffusion models are trained on a filtered version of the Objaverse dataset [7] with ∼ 50k 3D data. To enhance the quality and robustness, we design a series of strategies into our framework, including the noise offset channel in the multi-view diffusion training process to correct the discrepancy between training and inference [21], a stricter dataset filtering policy, and an expansion regularization to avoid normal collapse in mesh reconstruction. Overall, our method can generate high-fidelity, diverse, and multiview consistent meshes from single-view wild images within 30 seconds, as shown in Figure 1.
We conduct extensive experiments on various wild 2D images with different styles. The experiments verify the efficacy of our framework and show that our Unique3D outperforms existing methods for high fidelity, geometric details, high resolution, and strong generalizability.
In summary, our contributions are:
• We propose a novel image-to-3D framework called Unique3D that holistically archives a leading level of high-fidelity, efficiency, and generalizability among current methods.
• We introduce a multi-level upscale strategy to progressively generate higher-resolution RGB images with the corresponding normal maps.
• We design a novel instant and consistent mesh reconstruction algorithm (ISOMER) to reconstruct 3D meshes with intricate geometric details and texture from RGB images and normal maps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 513d47e5-933c-4da8-92dd-580e878467b4Cited by top-tier papers64
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and UnderstandingJunliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie et al.NeurIPS 2025 · 42 citations
- CAST: Component-Aligned 3D Scene Reconstruction from an RGB ImageKaixin Yao, Longwen Zhang, Xinhao Yan, Yan Zeng et al.SIGGRAPH 2025 · 30 citations
- ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via GenerationJiahao Chang, Chongjie Ye, Yushuang Wu, Yuantao Chen et al.ICLR 2026 · 30 citations
- Dimensionx: Create Any 3D and 4D Scenes From a Single Image With Decoupled Video DiffusionWenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen et al.ICCV 2025 · 11 citations
- Hi3dgen: High-Fidelity 3D Geometry Generation From Images Via Normal BridgingChongjie Ye, Yushuang Wu, Ziteng Lu, Jiahao Chang et al.ICCV 2025 · 11 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Wonder3D: Single Image to 3D Using Cross-Domain DiffusionXiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu et al.CVPR 2024 · 269 citations
- HumanRef: Single Image to 3D Human Generation via Reference-Guided DiffusionJingbo Zhang, Xiaoyu Li, Qi Zhang, Yanpei Cao et al.CVPR 2024 · 15 citations
- Animate3D: Animating Any 3D Model with Multi-view Video DiffusionYanqin Jiang, Chaohui Yu, Chenjie Cao, Fan Wang et al.NeurIPS 2024 · 65 citations
- UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided TexturesMingyuan Zhou, Rakib Hyder, Ziwei Xuan, Guojun QiCVPR 2024 · 8 citations
- ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent SynthesisXiangjun Gao, Xiaoyu Li, Chaopeng Zhang, Qi Zhang et al.CVPR 2024
