Generalized Deep 3D Shape Prior via Part-Discretized Diffusion Process
Yuhan Li, Yishun Dou, Xuanhong Chen, Bingbing Ni, Yilin Sun, Yutian Liu, Fuzhen Wang
Abstract
We develop a generalized 3D shape generation prior model, tailored for multiple 3D tasks including unconditional shape generation, point cloud completion, and cross-modality shape generation, etc. On one hand, to precisely capture local fine detailed shape information, a vector quantized variational autoencoder (VQ-VAE) is utilized to index local geometry from a compactly learned codebook based on a broad set of task training data. On the other hand, a discrete diffusion generator is introduced to model the inherent structural dependencies among different tokens. In the meantime, a multi-frequency fusion module (MFM) is developed to suppress high-frequency shape feature fluctuations, guided by multi-frequency contextual information. The above designs jointly equip our proposed 3D shape prior model with high-fidelity, diverse features as well as the capability of cross-modality alignment, and extensive experiments have demonstrated superior performances on various 3D shape generation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09609f22-af9d-4482-bf42-df8f3bb4f250Cited by top-tier papers10
- FocalDreamer: Text-Driven 3D Editing via Focal-Fusion AssemblyYuhan Li, Yishun Dou, Yue Shi, Yu Lei et al.AAAI 2024 · 91 citations
- MeshFormer : High-Quality Mesh Generation with 3D-Guided Reconstruction ModelMinghua Liu, Chong Zeng, Xinyue Wei, Ruoxi Shi et al.NeurIPS 2024 · 73 citations
- Neural Residual Diffusion Models for Deep Scalable Vision GenerationZhiyuan Ma, Liangliang Zhao, Biqing Qi, Bowen ZhouNeurIPS 2024 · 15 citations
- LAM3D: Large Image-Point Clouds Alignment Model for 3D Reconstruction from Single ImageRuikai Cui, Xibin Song, Weixuan Sun, Senbo Wang et al.NeurIPS 2024 · 7 citations
- GuideFlow3D: Optimization-Guided Rectified Flow For Appearance TransferSayan Deb Sarkar, Sinisa Stekovic, Vincent Lepetit, Iro ArmeniNeurIPS 2025 · 3 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationZibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng et al.NeurIPS 2023 · 279 citations
- GaussianAnything: Interactive Point Cloud Flow Matching for 3D GenerationYushi Lan, Shangchen Zhou, Zhaoyang Lyu, Fangzhou Hong et al.ICLR 2025
- Class-Partitioned VQ-VAE and Latent Flow Matching for Point Cloud Scene GenerationDasith de Silva Edirimuni, Ajmal Saeed MianAAAI 2026
- 3D Shape Generation and Completion through Point-Voxel DiffusionLinqi Zhou, Yilun Du, Jiajun WuICCV 2021 · 681 citations
- TAR3D: Creating High-Quality 3D Assets Via Next-Part PredictionXuying Zhang, Yutong Liu, Yangguang Li, Renrui Zhang et al.ICCV 2025 · 4 citations
