Sketch and Text Guided Diffusion Model for Colored Point Cloud Generation
Zijie Wu, Yaonan Wang, Mingtao Feng, He Xie, Ajmal Mian
摘要
Diffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently ambiguous and lack details. In this paper, we propose a sketch and text guided probabilistic diffusion model for colored point cloud generation that conditions the denoising process jointly with a hand drawn sketch of the object and its textual description. We incrementally diffuse the point coordinates and color values in a joint diffusion process to reach a Gaussian distribution. Colored point cloud generation thus amounts to learning the reverse diffusion process, conditioned by the sketch and text, to iteratively recover the desired shape and color. Specifically, to learn effective sketch-text embedding, our model adaptively aggregates the joint embedding of text prompt and the sketch based on a capsule attention network. Our model uses staged diffusion to generate the shape and then assign colors to different parts conditioned on the appearance prompt while preserving precise shapes from the first stage. This gives our model the flexibility to extend to multiple tasks, such as appearance re-editing and part segmentation. Experimental results demonstrate that our model outperforms recent state-of-the-art in point cloud generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Tactile DreamFusion: Exploiting Tactile Sensing for 3D GenerationRuihan Gao, Kangle Deng, Gengshan Yang, Wenzhen Yuan 等NeurIPS 2024 · 被引用 13 次
- VAPCNet: Viewpoint-Aware 3D Point Cloud CompletionZhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 等ICCV 2023 · 被引用 10 次
- Amodal3R: Amodal 3D Reconstruction from Occluded 2D ImagesTianhao Wu, Chuanxia Zheng, Frank Guan, Andrea Vedaldi 等ICCV 2025 · 被引用 9 次
- Modelling complex vector drawings with stroke-cloudsAlexander Ashcroft, Ayan Das, Yulia Gryaditskaya, Zhiyu Qu 等ICLR 2024 · 被引用 5 次
- A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene GenerationWentao Qu, Guofeng Mei, Yang Wu, Yongshun Gong 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume GenerationDelin An, Chaoli WangCVPR 2026
- Diffusion Probabilistic Models for 3D Point Cloud GenerationShitong Luo, Wei HuCVPR 2021
- SketchDream: Sketch-based Text-To-3D Generation and EditingFeng-Lin Liu, Hongbo Fu, Yu-Kun Lai, Lin GaoSIGGRAPH 2024 · 被引用 30 次
- 3D Shape Generation and Completion through Point-Voxel DiffusionLinqi Zhou, Yilun Du, Jiajun WuICCV 2021 · 被引用 681 次
- Control3D: Towards Controllable Text-to-3D GenerationYang Chen, Yingwei Pan, Yehao Li, Ting Yao 等ACM MM 2023 · 被引用 54 次
