Spice·E: Structural Priors in 3D Diffusion using Cross-Entity Attention
Etai Sella, Gal Fiebelman, Noam Atia, Hadar Averbuch-Elor
Abstract
We are witnessing rapid progress in automatically generating and manipulating 3D assets due to the availability of pretrained text-to-image diffusion models. However, time-consuming optimization procedures are required for synthesizing each sample, hindering their potential for democratizing 3D content creation. Conversely, 3D diffusion models now train on million-scale 3D datasets, yielding high-quality text-conditional 3D samples within seconds. In this work, we present Spice · E – a neural network that adds structural guidance to 3D diffusion models, extending their usage beyond text-conditional generation. At its core, our framework introduces a cross-entity attention mechanism that allows for multiple entities—in particular, paired input and guidance 3D shapes—to interact via their internal representations within the denoising network. We utilize this mechanism for learning task-specific structural priors in 3D diffusion models from auxiliary guidance shapes. We show that our approach supports a variety of applications, including 3D stylization, semantic shape editing and text-conditional abstraction-to-3D, which transforms primitive-based abstractions into highly-expressive shapes. Extensive experiments demonstrate that Spice · E achieves SOTA performance over these tasks while often being considerably faster than alternative methods. Importantly, this is accomplished without tailoring our approach for any specific task. We will release our code and trained models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04a994fa-97ed-4507-8dac-65be83f53804Cited by top-tier papers4
- GenPara: Enhancing the 3D Design Editing Process by Inferring Users' Regions of Interest with Text-Conditional Shape ParametersJiin Choi, Seung Won Lee, Kyung Hoon HyunCHI 2025 · 6 citations
- Blended Point Cloud Diffusion for Localized Text-Guided Shape EditingEtai Sella, Noam Atia, Ron Mokady, Hadar Averbuch-ElorICCV 2025 · 4 citations
- Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based AbstractionsEtai Sella, Hao Phung, Nitay Amiel, Or Litany et al.SIGGRAPH 2026 · 2 citations
- ShapeUP: Scalable Image-Conditioned 3D EditingInbar Gat, Dana Cohen-Bar, Guy Levy, Elad Richardson et al.SIGGRAPH 2026
Builds on47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- PI3D: Efficient Text-to-3D Generation with Pseudo-Image DiffusionYing-Tian Liu, Yuan-Chen Guo, Guan Luo, Heyi Sun et al.CVPR 2024
- Laconic: A 3D Layout Adapter for Controllable Image CreationLéopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks OvsjanikovICCV 2025
- Geometry Image Diffusion: Fast and Data-Efficient Text-to-3D with Image-Based Surface RepresentationSlava Elizarov, Ciara Rowles, Simon DonnéICLR 2025
- Vox-E: Text-guided Voxel Editing of 3D ObjectsEtai Sella, Gal Fiebelman, Peter Hedman, Hadar Averbuch-ElorICCV 2023 · 122 citations
- DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D GenerationYukun Huang, Jianan Wang, Yukai Shi, Boshi Tang et al.ICLR 2024 · 79 citations
