Taming Mode Collapse in Score Distillation for Text-to-3D Generation
Peihao Wang, Dejia Xu, Zhiwen Fan, Dilin Wang, Sreyas Mohan, Forrest N. Iandola, Rakesh Ranjan, Yilei Li, Qiang Liu, Zhangyang Wang, Vikas Chandra
Abstract
Despite the remarkable performance of score distillation in text-to-3D generation, such techniques notoriously suf-fer from view inconsistency issues, also known as “Janus” artifact, where the generated objects fake each view with multiple front faces. Although empirically effective methods have approached this problem via score debiasing or prompt engineering, a more rigorous perspective to explain and tackle this problem remains elusive. In this paper, we reveal that the existing score distillation-based text-to-3D generation frameworks degenerate to maximal likelihood seeking on each view independently and thus suffer from the mode collapse problem, manifesting as the Janus artifact in practice. To tame mode collapse, we improve score distillation by re-establishing the entropy term in the corresponding variational objective, which is applied to the distribution of rendered images. Maximizing the entropy encourages diversity among different views in generated 3D assets, thereby mitigating the Janus problem. Based on this new objective, we derive a new update rule for 3D score distillation, dubbed Entropic Score Distillation (ESD). We theoretically reveal that ESD can be simplified and implemented by just adopting the classifier-free guidance trick upon variational score distillation. Although embarrassingly straightforward, our extensive experiments demonstrate that ESD can be an effective treatment for Janus artifacts in score distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f686ebc7-484e-4cbf-b08d-4a2e39becc9eCited by top-tier papers22
- One-Step Effective Diffusion Network for Real-World Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhiyuan Ma, Lei ZhangNeurIPS 2024 · 319 citations
- Score Distillation via Reparametrized DDIMArtem Lukoianov, Haitz Sáez de Ocáriz Borde, Kristjan H. Greenewald, Vitor Guizilini et al.NeurIPS 2024 · 50 citations
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language ModelsSilin Cheng, Kai HanNeurIPS 2025 · 7 citations
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry GroundingJulian Ost, Andrea Ramazzina, Amogh Joshi, Maximilian Bömer et al.AAAI 2026 · 6 citations
- LCGen: Mining in Low-Certainty Generation for View-consistent Text-to-3DZeng Tao, Tong Yang, Junxiong Lin, Xinji Mai et al.NeurIPS 2024 · 4 citations
Builds on34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- Debiasing Scores and Prompts of 2D Diffusion for View-consistent Text-to-3D GenerationSusung Hong, Donghoon Ahn, Seungryong KimNeurIPS 2023 · 46 citations
- RecDreamer: Consistent Text-to-3D Generation via Uniform Score DistillationChenxi Zheng, Yihong Lin, Bangzhen Liu, Xuemiao Xu et al.ICLR 2025
- Text-to-3D with Classifier Score DistillationXin Yu, Yuan-Chen Guo, Yangguang Li, Ding Liang et al.ICLR 2024 · 132 citations
- Portrait3D: Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs PriorYiqian Wu, Hao Xu, Xiangjun Tang, Xien Chen et al.SIGGRAPH 2024 · 14 citations
- Consistent Flow Distillation for Text-to-3D GenerationRunjie Yan, Yinbo Chen, Xiaolong WangICLR 2025
