Apply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D Generation
Yiming Qin, Zhu Xu, Yang Liu
Abstract
Input Text: A man in black coat, yellow shirt, pink trousers, blue shoes and green hat is waving black Cross-attention map A man in black coat, yellow shirt, pink trousers, blue shoes and green hat is waving (a) Visualization of cross-attention map.
For longer text prompts, 2D Stable Diffusion (SD) [26] fails to accurately associate the word "black" with the correct spatial location in the generated image. This limitation poses a challenge for methods [15,27] lifting 2D to 3D using SD effectively.
Recent text-to-3D generation models have demonstrated remarkable abilities in producing high-quality 3D assets. Despite their great advancements, current models struggle to generate satisfying 3D objects with complex attributes. The difficulty for such complex attributes 3D generation arises from two aspects: (1) existing text-to-3D approaches typi- * Corresponding author.
cally lift text-to-image models to extract semantics via text encoders, while the text encoder exhibits limited comprehension ability for long descriptions, leading to deviated cross-attention focus, subsequently wrong attribute binding in generated results. (2) Objects with complex attributes often exhibit occlusion relationships between different parts, which demands a reasonable generation order as well as explicit disentanglement of different parts to enable structural coherent and attribute following results. Though some This CVPR paper is the Open Access version, provided by the Computer Vision Foundation.
Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Interact-Custom: Customized Human Object Interaction Image GenerationZhu Xu, Zhaowen Wang, Yuxin Peng, Yang LiuACM MM 2025 · 1 citation
- Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without TrainingHexiao Lu, Xiaokun Sun, Zeyu Cai, Hao Guo et al.CVPR 2026
- Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMsWentao Mo, Yang LiuICML 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 1 citation
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
- CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion ModelsTuna Han Salih Meral, Enis Simsar, Federico Tombari, Pinar YanardagCVPR 2024 · 11 citations
- What the DAAM: Interpreting Stable Diffusion Using Cross AttentionRaphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang et al.ACL 2023 · 93 citations
- Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsXinhua Cheng, Tianyu Yang, Jianan Wang, Yu Li et al.ICLR 2024 · 58 citations
