Sketch and Text Guided Diffusion Model for Colored Point Cloud Generation
Zijie Wu, Yaonan Wang, Mingtao Feng, He Xie, Ajmal Mian
Abstract
Diffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently ambiguous and lack details. In this paper, we propose a sketch and text guided probabilistic diffusion model for colored point cloud generation that conditions the denoising process jointly with a hand drawn sketch of the object and its textual description. We incrementally diffuse the point coordinates and color values in a joint diffusion process to reach a Gaussian distribution. Colored point cloud generation thus amounts to learning the reverse diffusion process, conditioned by the sketch and text, to iteratively recover the desired shape and color. Specifically, to learn effective sketch-text embedding, our model adaptively aggregates the joint embedding of text prompt and the sketch based on a capsule attention network. Our model uses staged diffusion to generate the shape and then assign colors to different parts conditioned on the appearance prompt while preserving precise shapes from the first stage. This gives our model the flexibility to extend to multiple tasks, such as appearance re-editing and part segmentation. Experimental results demonstrate that our model outperforms recent state-of-the-art in point cloud generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ca99d40-a82b-405d-a0b5-6cffb091349fCited by top-tier papers18
- Tactile DreamFusion: Exploiting Tactile Sensing for 3D GenerationRuihan Gao, Kangle Deng, Gengshan Yang, Wenzhen Yuan et al.NeurIPS 2024 · 13 citations
- VAPCNet: Viewpoint-Aware 3D Point Cloud CompletionZhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang et al.ICCV 2023 · 10 citations
- Amodal3R: Amodal 3D Reconstruction from Occluded 2D ImagesTianhao Wu, Chuanxia Zheng, Frank Guan, Andrea Vedaldi et al.ICCV 2025 · 9 citations
- Modelling complex vector drawings with stroke-cloudsAlexander Ashcroft, Ayan Das, Yulia Gryaditskaya, Zhiyu Qu et al.ICLR 2024 · 5 citations
- A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene GenerationWentao Qu, Guofeng Mei, Yang Wu, Yongshun Gong et al.CVPR 2026 · 4 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume GenerationDelin An, Chaoli WangCVPR 2026
- Diffusion Probabilistic Models for 3D Point Cloud GenerationShitong Luo, Wei HuCVPR 2021
- SketchDream: Sketch-based Text-To-3D Generation and EditingFeng-Lin Liu, Hongbo Fu, Yu-Kun Lai, Lin GaoSIGGRAPH 2024 · 30 citations
- 3D Shape Generation and Completion through Point-Voxel DiffusionLinqi Zhou, Yilun Du, Jiajun WuICCV 2021 · 681 citations
- Control3D: Towards Controllable Text-to-3D GenerationYang Chen, Yingwei Pan, Yehao Li, Ting Yao et al.ACM MM 2023 · 54 citations
