HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape Generation
Zhiying Leng, Tolga Birdal, Xiaohui Liang, Federico Tombari
Abstract
3D shape generation from text is a fundamental task in 3D representation learning. The text-shape pairs exhibit a hierarchical structure, where a general text like "chair" covers all 3D shapes of the chair, while more detailed prompts refer to more specific shapes. Furthermore, both text and 3D shapes are inherently hierarchical structures. However, existing Text2Shape methods, such as SDFusion, do not exploit that. In this work, we propose HyperSD-Fusion, a dual-branch diffusion model that generates 3D shapes from a given text. Since hyperbolic space is suitable for handling hierarchical data, we propose to learn the hierarchical representations of text and 3D shapes in hyperbolic space. First, we introduce a hyperbolic text-image encoder to learn the sequential and multi-modal hierarchical features of text in hyperbolic space. In addition, we design a hyperbolic text-graph convolution module to learn the hierarchical features of text in hyperbolic space. In order to fully utilize these text features, we introduce a dual-branch structure to embed text features in 3D feature space. At last, to endow the generated 3D shapes with a hierarchical structure, we devise a hyperbolic hierarchical loss. Our method is the first to explore the hyperbolic hierarchical representation for text-to-shape generation. Experimental results on the existing text-to-shape paired dataset, Text2Shape, achieved state-of-the-art results. We release our implementation under HyperSDFusion.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 155fdc2c-82a2-4151-af22-892235157febCited by top-tier papers8
- Generative Fractional Diffusion ModelsGabriel Nobis, Maximilian Springenberg, Marco Aversa, Michael Detzel et al.NeurIPS 2024 · 19 citations
- HYPDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-Shot Image GenerationLingxiao Li, Kaixuan Fan, Boqing Gong, Xiangyu YueICCV 2025 · 5 citations
- Fractional Diffusion Bridge ModelsGabriel Nobis, Maximilian Springenberg, Arina Belova, Rembert Daems et al.NeurIPS 2025 · 4 citations
- Intrinsic Lorentz Neural NetworkXianglong Shi, Ziheng Chen, Yunhan Jiang, Nicu SebeICLR 2026 · 3 citations
- Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual RepresentationBoyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.CVPR 2026 · 2 citations
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Learning Hierarchical Hyperbolic Mixture Model for Part-aware 3D GenerationQitong Yang, Mingtao Feng, Zijie Wu, Huixin Zhu et al.CVPR 2026
- Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic SpaceHuan Kang, Hui Li, Tianyang Xu, Tao Zhou et al.ICML 2026
- Diffusion-SDF: Text-to-Shape via Voxelized DiffusionMuheng Li, Yueqi Duan, Jie Zhou, Jiwen LuCVPR 2023
- Hyperbolic Graph Diffusion ModelLingfeng Wen, Xuan Tang, Mingjie Ouyang, Xiangxiang Shen et al.AAAI 2024 · 16 citations
- Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D RetrievalWenrui Li, Yidan Lu, Yeyu Chai, Rui Zhao et al.AAAI 2026
