Geometry Image Diffusion: Fast and Data-Efficient Text-to-3D with Image-Based Surface Representation
Slava Elizarov, Ciara Rowles, Simon Donné
摘要
Generating high-quality 3D objects from textual descriptions remains a challenging problem due to computational cost, the scarcity of 3D data, and complex 3D representations. We introduce Geometry Image Diffusion (GIMDiffusion), a novel Text-to-3D model that utilizes geometry images to efficiently represent 3D shapes using 2D images, thereby avoiding the need for complex 3D-aware architectures. By integrating a Collaborative Control mechanism, we exploit the rich 2D priors of existing Text-to-Image models such as Stable Diffusion. This enables strong generalization even with limited 3D training data (allowing us to use only high-quality training data) as well as retaining compatibility with guidance techniques such as IPAdapter. In short, GIMDiffusion enables the generation of 3D assets at speeds comparable to current Text-to-Image models. The generated objects consist of semantically meaningful, separate parts and include internal structures, enhancing both usability and versatility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Generative Human Geometry DistributionXiangjun Tang, Biao Zhang, Peter WonkaICLR 2026 · 被引用 4 次
- RomanTex: Decoupling 3D-Aware Rotary Positional Embedded Multi-Attention Network for Texture SynthesisYifei Feng, Mingxin Yang, Shuhui Yang, Sheng Zhang 等ICCV 2025 · 被引用 3 次
- Repurposing 2D Diffusion Models with Gaussian Atlas for 3D GenerationTiange Xiang, Kai Li, Chengjiang Long, Christian Häne 等ICCV 2025 · 被引用 1 次
- Garment Particles: A 2D-3D Symmetric Garment Representation for Generation and EditingKiyohiro Nakayama, I-Chao Shen, Ruofan Liu, Yiming Wang 等SIGGRAPH 2026 · 被引用 1 次
- SwiftTailor: Efficient 3D Garment Generation with Geometry Image RepresentationPhuc Pham, Uy Dieu Tran, Binh-Son Hua, Phong NguyenCVPR 2026
它引用的顶会 Paper29
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- Laconic: A 3D Layout Adapter for Controllable Image CreationLéopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks OvsjanikovICCV 2025
- PI3D: Efficient Text-to-3D Generation with Pseudo-Image DiffusionYing-Tian Liu, Yuan-Chen Guo, Guan Luo, Heyi Sun 等CVPR 2024
- TexFusion: Synthesizing 3D Textures with Text-Guided Image Diffusion ModelsTianshi Cao, Karsten Kreis, Sanja Fidler, Nicholas Sharp 等ICCV 2023 · 被引用 103 次
- Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelJiahao Li, Hao Tan, Kai Zhang, Zexiang Xu 等ICLR 2024 · 被引用 408 次
- I2V3D: Controllable Image-to-Video Generation with 3D GuidanceZhiyuan Zhang, Dongdong Chen, Jing LiaoICCV 2025 · 被引用 3 次
