Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation
Zibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng, Rui Wang, Pei Cheng, Bin Fu, Tao Chen, Gang Yu, Shenghua Gao
摘要
We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone to producing inconsistent results with the conditions because 3D shapes have an additional dimension whose distribution significantly differs from that of 2D images and texts. To bridge the domain gap among the three modalities and facilitate multi-modal-conditioned 3D shape generation, we explore representing 3D shapes in a shape-image-text-aligned space. Our framework comprises two models: a Shape-Image-Text-Aligned Variational Auto-Encoder (SITA-VAE) and a conditional Aligned Shape Latent Diffusion Model (ASLDM). The former model encodes the 3D shapes into the shape latent space aligned to the image and text and reconstructs the fine-grained 3D neural fields corresponding to given shape embeddings via the transformer-based decoder. The latter model learns a probabilistic mapping function from the image or text space to the latent shape space. Our extensive experiments demonstrate that our proposed approach can generate higher-quality and more diverse 3D shapes that better semantically conform to the visual or textural conditional inputs, validating the effectiveness of the shape-image-text-aligned space for cross-modality 3D shape generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper100
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang 等NeurIPS 2024 · 被引用 251 次
- Native and Compact Structured Latents for 3D GenerationJianfeng Xiang, Xiaoxue Chen, Sicheng Xu, Ruicheng Wang 等CVPR 2026 · 被引用 177 次
- CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D AssetsLongwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu 等SIGGRAPH 2024 · 被引用 148 次
- MeshXL: Neural Coordinate Field for Generative 3D Foundation ModelsSijin Chen, Xin Chen, Anqi Pang, Xianfang Zeng 等NeurIPS 2024 · 被引用 125 次
- Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionShuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng 等NeurIPS 2025 · 被引用 114 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- 3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion ModelsBiao Zhang, Jiapeng Tang, Matthias Nießner, Peter WonkaSIGGRAPH 2023 · 被引用 172 次
- A3D: Does Diffusion Dream about 3D Alignment?Savva Victorovich Ignatyev, Nina Konovalova, Daniil Selikhanovych, Oleg Voynov 等ICLR 2025
- Text-Conditional Attribute Alignment Across Latent Spaces for 3D Controllable Face Image SynthesisFeifan Xu, Rui Li, Si Wu, Yong Xu 等CVPR 2024
- Generalized Deep 3D Shape Prior via Part-Discretized Diffusion ProcessYuhan Li, Yishun Dou, Xuanhong Chen, Bingbing Ni 等CVPR 2023
- Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion ModelingYanchen Luo, Zhiyuan Liu, Yi Zhao, Sihang Li 等NeurIPS 2025 · 被引用 11 次
