AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
Jiayin Zhu, Linlin Yang, Yicong Li, Angela Yao
Abstract
Optimization‐based text‑to‑3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent trajectories that suppress or merge semantic cues, leading to "semantic over-smoothing" artifacts. As such, we reformulate text‑to‑3D optimization as mapping a dynamically evolving source distribution to a fixed target distribution. We cast the problem into a dual‑conditioned latent space, conditioned on both the text prompt and the intermediately rendered image. Given this joint setup, we observe that the image condition naturally anchors the current source distribution. Building on this insight, we introduce AnchorDS, an improved score distillation mechanism that provides state‑anchored guidance with image conditions and stabilizes generation. We further penalize erroneous source estimates and design a lightweight filter strategy and fine‑tuning strategy that refines the anchor with negligible overhead. AnchorDS produces finer-grained detail, more natural colours, and stronger semantic consistency, particularly for complex prompts, while maintaining efficiency. Extensive experiments show that our method surpasses previous methods in both quality and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fb851e2-c8de-43f8-89b9-9faba924f22dCited by top-tier papers1
Ask how each one uses itBuilds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai et al.ICLR 2024 · 973 citations
Related papers
- Target-Balanced Score DistillationZhou Xu, Qi Wang, Yuxiao Yang, Luyuan Zhang et al.AAAI 2026
- Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling PriorZike Wu, Pan Zhou, Xuanyu Yi, Xiaoding Yuan et al.CVPR 2024 · 15 citations
- Rethinking Score Distillation as a Bridge Between Image DistributionsDavid McAllister, Songwei Ge, Jia-Bin Huang, David Jacobs et al.NeurIPS 2024 · 43 citations
- Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D GenerationZiying Li, Xuequan Lu, Xinkui Zhao, Guanjie Cheng et al.NeurIPS 2025 · 4 citations
- Stable Score DistillationHaiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen et al.ICCV 2025 · 2 citations
