Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion
Zhenwei Wang, Tengfei Wang, Zexin He, Gerhard Petrus Hancke, Ziwei Liu, Rynson W. H. Lau
Abstract
In 3D modeling, designers often use an existing 3D model as a reference to create new ones. This practice has inspired the development of Phidias, a novel generative model that uses diffusion for reference-augmented 3D generation. Given an image, our method leverages a retrieved or user-provided 3D reference model to guide the generation process, thereby enhancing the generation quality, generalization ability, and controllability. Our model integrates three key components: 1) meta-ControlNet that dynamically modulates the conditioning strength, 2) dynamic reference routing that mitigates misalignment between the input image and 3D reference, and 3) self-reference augmentations that enable self-supervised training with a progressive curriculum. Collectively, these designs result in a clear improvement over existing methods. Phidias establishes a unified framework for 3D generation using text, image, and 3D conditions with versatile applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative ModelingBowen Zhang, Yiji Cheng, Jiaolong Yang, Chunyu Wang et al.NeurIPS 2024 · 49 citations
- SpaceControl: Introducing Test-Time Spatial Control to 3D Generative ModelingElisabetta Fedele, Francis Engelmann, Ian Huang, Or Litany et al.ICLR 2026 · 11 citations
- UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted OptimizationQianfeng Yang, Qiyuan Guan, Xiang Chen, Jiyu Jin et al.CVPR 2026 · 6 citations
- Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view DiffusionZeren Xiong, Zikun Chen, Zedong Zhang, Xiang Li et al.ACM MM 2025 · 2 citations
- RefAny3D: 3D Asset-Referenced Diffusion Models for Image GenerationHanzhuo Huang, Qingyang Bao, Zekai Gu, Zhongshuo Du et al.ICLR 2026 · 1 citation
Builds on20
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi et al.ICLR 2024 · 813 citations
Related papers
- Control3D: Towards Controllable Text-to-3D GenerationYang Chen, Yingwei Pan, Yehao Li, Ting Yao et al.ACM MM 2023 · 54 citations
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationZilong Chen, Yikai Wang, Wenqiang Sun, Feng Wang et al.CVPR 2025
- CMD: Controllable Multiview Diffusion for 3D Editing and Progressive GenerationPeng Li, Suizhi Ma, Jialiang Chen, Yuan Liu et al.SIGGRAPH 2025 · 8 citations
- SketchDream: Sketch-based Text-To-3D Generation and EditingFeng-Lin Liu, Hongbo Fu, Yu-Kun Lai, Lin GaoSIGGRAPH 2024 · 30 citations
- Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D PriorCheng Chen, Xiaofeng Yang, Fan Yang, Chengzeng Feng et al.CVPR 2024
