Text-to-3D Generation with Bidirectional Diffusion Using Both 2D and 3D Priors
Lihe Ding, Shaocong Dong, Zhanpeng Huang, Zibin Wang, Yiyuan Zhang, Kaixiong Gong, Dan Xu, Tianfan Xue
Abstract
≈ 20min A yellow and green oil painting style eagle head (c) ProlificDreamer (b) Zero-123 (a) Shap-E Figure 1. Our BiDiff can efficiently generate high-quality 3D objects. It alleviates all these issues in previous 3D generative models: (a) low-texture quality, (b) multi-view inconsistency, and (c) geometric incorrectness (e.g., multi-face Janus problem). The outputs of our model can be further combined with optimization-based methods (e.g., ProlificDreamer) to generate better 3D geometries with slightly longer processing time (bottom row).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- FullPart: Generating each 3D Part at Full ResolutionLihe Ding, Shaocong Dong, Yaokun Li, Chenjian Gao et al.ICLR 2026 · 17 citations
- From One to More: Contextual Part Latents for 3D GenerationShaocong Dong, Lihe Ding, Xiao Chen, Yaokun Li et al.ICCV 2025 · 4 citations
- MaterialMVP: Illumination-Invariant Material Generation via Multi-View PBR DiffusionZebin He, Mingxin Yang, Shuhui Yang, Yixuan Tang et al.ICCV 2025 · 3 citations
- DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video GenerationDonglin Di, He Feng, Wenzhang Sun, Yongjia Ma et al.ICCV 2025 · 1 citation
- A General Framework to Boost 3D GS Initialization for Text-to-3D Generation by Lexical RichnessLutao Jiang, Hangyu Li, Lin WangACM MM 2024 · 1 citation
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained DiffusionZehuan Huang, Hao Wen, Junting Dong, Yaohui Wang et al.CVPR 2024
- Enhancing 3D Fidelity of Text-to-3D using Cross-View CorrespondencesSeungwook Kim, Kejie Li, Xueqing Deng, Yichun Shi et al.CVPR 2024
- Instant3dit: Multiview Inpainting for Fast Editing of 3D ObjectsAmir Barda, Matheus Gadelha, Vladimir G. Kim, Noam Aigerman et al.CVPR 2025
- Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play DeformationQiao Yu, Xianzhi Li, Yuan Tang, Xu Han et al.CVPR 2025
- Bolt3D: Generating 3D Scenes in SecondsStanislaw Szymanowicz, Jason Y. Zhang, Pratul P. Srinivasan, Ruiqi Gao et al.ICCV 2025 · 11 citations
