Text-to-3D with Classifier Score Distillation
Xin Yu, Yuan-Chen Guo, Yangguang Li, Ding Liang, Song-Hai Zhang, Xiaojuan Qi
Abstract
Text-to-3D generation has made remarkable progress recently, particularly with methods based on Score Distillation Sampling (SDS) that leverages pre-trained 2D diffusion models. While the usage of classifier-free guidance is well acknowledged to be crucial for successful optimization, it is considered an auxiliary trick rather than the most essential component. In this paper, we re-evaluate the role of classifier-free guidance in score distillation and discover a surprising finding: the guidance alone is enough for effective text-to-3D generation tasks. We name this method Classifier Score Distillation (CSD), which can be interpreted as using an implicit classification model for generation. This new perspective reveals new insights for understanding existing techniques. We validate the effectiveness of CSD across a variety of text-to-3D tasks including shape generation, texture synthesis, and shape editing, achieving results superior to those of state-of-the-art methods. Our project page is https://xinyu-andy.github.io/Classifier-Score-Distillation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers58
- L4GM: Large 4D Gaussian Reconstruction ModelJiawei Ren, Cheng Xie, Ashkan Mirzaei, Hanxue Liang et al.NeurIPS 2024 · 173 citations
- SCube: Instant Large-Scale Scene Reconstruction using VoxSplatsXuanchi Ren, Yifan Lu, Hanxue Liang, Jay Zhangjie Wu et al.NeurIPS 2024 · 63 citations
- Score Distillation via Reparametrized DDIMArtem Lukoianov, Haitz Sáez de Ocáriz Borde, Kristjan H. Greenewald, Vitor Guizilini et al.NeurIPS 2024 · 50 citations
- HumanGaussian: Text-Driven 3D Human Generation with Gaussian SplattingXian Liu, Xiaohang Zhan, Jiaxiang Tang, Ying Shan et al.CVPR 2024 · 42 citations
- GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian GenerationChubin Zhang, Hongliang Song, Yi Wei, Chen Yu et al.NeurIPS 2024 · 40 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Noise-free Score DistillationOren Katzir, Or Patashnik, Daniel Cohen-Or, Dani LischinskiICLR 2024 · 101 citations
- Stable Score DistillationHaiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen et al.ICCV 2025 · 2 citations
- VP3D: Unleashing 2D Visual Prompt for Text-to-3D GenerationYang Chen, Yingwei Pan, Haibo Yang, Ting Yao et al.CVPR 2024
- Rethinking Score Distilling Sampling for 3D Editing and GenerationXingyu Miao, Haoran Duan, Yang Long, Jungong HanICML 2025
- Teefusion: Blending Text Embeddings to Distill Classifier-Free GuidanceMinghao Fu, Guo-Hua Wang, Xiaohao Chen, Qing-Guo Chen et al.ICCV 2025
