VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation
Yang Chen, Yingwei Pan, Haibo Yang, Ting Yao, Tao Mei
摘要
Recent innovations on text-to-3D generation have featured Score Distillation Sampling (SDS), which enables the zero-shot learning of implicit 3D models (NeRF) by directly distilling prior knowledge from 2D diffusion models. However, current SDS-based models still struggle with intricate text prompts and commonly result in distorted 3D models with unrealistic textures or cross-view inconsistency issues. In this work, we introduce a novel Visual Prompt-guided text-to-3D diffusion model (VP3D) that explicitly unleashes the visual appearance knowledge in 2D visual prompt to boost text-to-3D generation. Instead of solely supervising SDS with text prompt, VP3D first capitalizes on 2D diffusion model to generate a high-quality image from input text, which subsequently acts as visual prompt to strengthen SDS optimization with explicit visual appearance. Meanwhile, we couple the SDS optimization with additional differentiable reward function that encourages rendering images of 3D models to better visually align with 2D visual prompt and semantically match with text prompt. Through extensive experiments, we show that the 2D Visual Prompt in our VP3D significantly eases the learning of visual appearance of 3D models and thus leads to higher visual fidelity with more detailed textures. It is also appealing in view that when replacing the self-generating visual prompt with a given reference image, VP3D is able to trigger a new task of stylized text-to-3D generation. Our project page is available at https://vp3d-cvpr24.github.io .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion ModelsHaibo Yang, Yang Chen, Yingwei Pan, Ting Yao 等ACM MM 2024 · 被引用 22 次
- Training-Free and Adaptive Sparse Attention for Efficient Long Video GenerationYifei Xia, Suhan Ling, Fangcheng Fu, Yujie Wang 等ICCV 2025 · 被引用 6 次
- FreeEnhance: Tuning-Free Image Enhancement via Content-Consistent Noising-and-Denoising ProcessYang Luo, Yiheng Zhang, Zhaofan Qiu, Ting Yao 等ACM MM 2024 · 被引用 4 次
- Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-To-3D GenerationYujie Zhang, Bingyang Cui, Qi Yang, Zhu Li 等ICCV 2025 · 被引用 3 次
- A General Framework to Boost 3D GS Initialization for Text-to-3D Generation by Lexical RichnessLutao Jiang, Hangyu Li, Lin WangACM MM 2024 · 被引用 1 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 等NeurIPS 2023 · 被引用 1,498 次
- Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D GenerationJunyoung Seo, Wooseok Jang, Minseop Kwak, Inès Hyeonsu Kim 等ICLR 2024 · 被引用 157 次
- Vox-E: Text-guided Voxel Editing of 3D ObjectsEtai Sella, Gal Fiebelman, Peter Hedman, Hadar Averbuch-ElorICCV 2023 · 被引用 122 次
- Target-Balanced Score DistillationZhou Xu, Qi Wang, Yuxiao Yang, Luyuan Zhang 等AAAI 2026
- Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D GenerationChaohui Yu, Qiang Zhou, Jingliang Li, Zhe Zhang 等ACM MM 2023 · 被引用 26 次
