Geometry Image Diffusion: Fast and Data-Efficient Text-to-3D with Image-Based Surface Representation
Slava Elizarov, Ciara Rowles, Simon Donné
Abstract
Generating high-quality 3D objects from textual descriptions remains a challenging problem due to computational cost, the scarcity of 3D data, and complex 3D representations. We introduce Geometry Image Diffusion (GIMDiffusion), a novel Text-to-3D model that utilizes geometry images to efficiently represent 3D shapes using 2D images, thereby avoiding the need for complex 3D-aware architectures. By integrating a Collaborative Control mechanism, we exploit the rich 2D priors of existing Text-to-Image models such as Stable Diffusion. This enables strong generalization even with limited 3D training data (allowing us to use only high-quality training data) as well as retaining compatibility with guidance techniques such as IPAdapter. In short, GIMDiffusion enables the generation of 3D assets at speeds comparable to current Text-to-Image models. The generated objects consist of semantically meaningful, separate parts and include internal structures, enhancing both usability and versatility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ad36331-4cf9-43b3-8765-23d85cfc905cCited by top-tier papers8
- Generative Human Geometry DistributionXiangjun Tang, Biao Zhang, Peter WonkaICLR 2026 · 4 citations
- RomanTex: Decoupling 3D-Aware Rotary Positional Embedded Multi-Attention Network for Texture SynthesisYifei Feng, Mingxin Yang, Shuhui Yang, Sheng Zhang et al.ICCV 2025 · 3 citations
- Repurposing 2D Diffusion Models with Gaussian Atlas for 3D GenerationTiange Xiang, Kai Li, Chengjiang Long, Christian Häne et al.ICCV 2025 · 1 citation
- Garment Particles: A 2D-3D Symmetric Garment Representation for Generation and EditingKiyohiro Nakayama, I-Chao Shen, Ruofan Liu, Yiming Wang et al.SIGGRAPH 2026 · 1 citation
- SwiftTailor: Efficient 3D Garment Generation with Geometry Image RepresentationPhuc Pham, Uy Dieu Tran, Binh-Son Hua, Phong NguyenCVPR 2026
Builds on29
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Laconic: A 3D Layout Adapter for Controllable Image CreationLéopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks OvsjanikovICCV 2025
- PI3D: Efficient Text-to-3D Generation with Pseudo-Image DiffusionYing-Tian Liu, Yuan-Chen Guo, Guan Luo, Heyi Sun et al.CVPR 2024
- TexFusion: Synthesizing 3D Textures with Text-Guided Image Diffusion ModelsTianshi Cao, Karsten Kreis, Sanja Fidler, Nicholas Sharp et al.ICCV 2023 · 103 citations
- Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelJiahao Li, Hao Tan, Kai Zhang, Zexiang Xu et al.ICLR 2024 · 408 citations
- I2V3D: Controllable Image-to-Video Generation with 3D GuidanceZhiyuan Zhang, Dongdong Chen, Jing LiaoICCV 2025 · 3 citations
