Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generation
Delin An, Chaoli Wang
摘要
Diffusion probabilistic models have demonstrated significant potential in generating high-quality, realistic medical images, providing a promising solution to the persistent challenge of data scarcity in the medical field. Nevertheless, producing 3D medical volumes with anatomically consistent structures under multimodal conditions remains a complex and unresolved problem. We introduce Sketch2CT, a multimodal diffusion framework for structure-aware 3D medical volume generation, jointly guided by a user-provided 2D sketch and a textual description that captures 3D geometric semantics. The framework initially generates 3D segmentation masks of the target organ from random noise, conditioned on both modalities. To effectively align and fuse these inputs, we propose two key modules that refine sketch features with localized textual cues and integrate global sketch-text representations. Built upon a capsule-attention backbone, these modules leverage the complementary strengths of sketches and text to produce anatomically accurate organ shapes. The synthesized segmentation masks subsequently guide a latent diffusion model for 3D CT volume synthesis, enabling realistic reconstruction of organ appearances that are consistent with user-defined sketches and descriptions. Extensive experiments on public CT datasets demonstrate that Sketch2CT achieves superior performance in generating multimodal medical volumes. Its controllable, low-cost generation pipeline enables principled, efficient augmentation of medical datasets. Code is available at https://github.com/adlsn/Sketch2CT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Big Self-Supervised Models Advance Medical Image ClassificationShekoofeh Azizi, Basil Mustafa, Fiona Ryan, Zachary Beaver 等ICCV 2021 · 被引用 695 次
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang 等ICCV 2019 · 被引用 218 次
相关 Paper
- Sketch and Text Guided Diffusion Model for Colored Point Cloud GenerationZijie Wu, Yaonan Wang, Mingtao Feng, He Xie 等ICCV 2023 · 被引用 55 次
- GuideGen: A Text-Guided Framework for Paired Full-torso Anatomy and CT Volume GenerationLinrui Dai, Rongzhao Zhang, Yongrui Yu, Xiaofan ZhangAAAI 2026
- PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask ConditionsMahesh Bhosale, Abdul Wasi, Yuanhao Zhai, Yunjie Tian 等ICCV 2025 · 被引用 8 次
- SketchDream: Sketch-based Text-To-3D Generation and EditingFeng-Lin Liu, Hongbo Fu, Yu-Kun Lai, Lin GaoSIGGRAPH 2024 · 被引用 30 次
- Locally Attentional SDF Diffusion for Controllable 3D Shape GenerationXin-Yang Zheng, Hao Pan, Peng-Shuai Wang, Xin Tong 等SIGGRAPH 2023 · 被引用 122 次
