MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual Invariant
Chenlu Zhan, Yu Lin, Gaoang Wang, Hongwei Wang, Jian Wu
摘要
Medical generative models, acknowledged for their high-quality sample generation ability, have accelerated the fast growth of medical applications. However, recent works concentrate on separate medical generation models for dis-tinct medical tasks and are restricted to inadequate medi-cal multimodal knowledge, constraining medical compre-hensive diagnosis. In this paper, we propose MedM2G, a Medical Multi-Modal Generative framework, with the key innovation to align, extract, and generate medical multimodal within a unified model. Extending beyond single or two medical modalities, we efficiently align medical multimodal through the central alignment approach in the unified space. Significantly, our framework extracts valuable clini-cal knowledge by preserving the medical visual invariant of each imaging modal, thereby enhancing specific medical information for multimodal generation. By conditioning the adaptive cross-guided parameters into the multi-flow diffusion framework, our model promotes flexible interactions among medical multimodalfor generation. MedM2G is the first medical generative model that unifies medical generation tasks of text-to-image, image-to-text, and unified generation of medical modalities (CT, MRI, X-ray). It performs 5 medical generation tasks across 10 datasets, consistently outperforming various state-of-the-art works.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant RepresentationZhaorui Tan, Xi Yang, Tan Pan, Tianyi Liu 等ICCV 2025 · 被引用 5 次
- MRGen: Segmentation Data Engine for Underrepresented MRI ModalitiesHaoning Wu, Ziheng Zhao, Ya Zhang, Yanfeng Wang 等ICCV 2025 · 被引用 3 次
- DAMPER: A Dual-Stage Medical Report Generation Framework with Coarse-Grained MeSH Alignment and Fine-Grained Hypergraph MatchingXiaofei Huang, Wenting Chen, Jie Liu, Qisheng Lu 等AAAI 2025 · 被引用 2 次
- Predicting Spatial Transcriptomics from Histology Images via High-Order Multi-Cell Interaction ModelingYouhan Sun, Jiahua Rao, Kangrui Du, Jiancong Xie 等CVPR 2026
- MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual RepresentationsZiyang Zhang, Yang Yu, Yucheng Chen, Xulei Yang 等CVPR 2025
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-AnalysisJunzhi Ning, Wei Li, Cheng Tang, Jiashi Lin 等ICML 2026 · 被引用 13 次
- SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task AlignmentWeiren Zhao, DONG Yi, Cheng ChenICML 2026
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu 等ICCV 2023 · 被引用 47 次
- Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion ModelQingyu Shi, Jinbin Bai, Zhuoran Zhao, Wenhao Chai 等ICLR 2026 · 被引用 40 次
- Rethinking Diffusion Bridge Model with Dual Alignments for Medical Image SynthesisJinbao Wei, Yuhang Chen, Zhijie Wang, Gang Yang 等ACM MM 2025 · 被引用 3 次
