Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert Feedback
Janet Wang, Yunbei Zhang, Zhengming Ding, Jihun Hamm
摘要
Paucity of medical data severely limits the generalizability of diagnostic ML models, as the full spectrum of disease variability can not be represented by a small clinical dataset. To address this, diffusion models (DMs) have been considered as a promising avenue for synthetic image generation and augmentation. However, they frequently produce medically inaccurate images, deteriorating the model performance. Expert domain knowledge is critical for synthesizing images that correctly encode clinical information, especially when data is scarce and quality outweighs quantity. Existing approaches for incorporating human feedback, such as reinforcement learning (RL) and Direct Preference Optimization (DPO), rely on robust reward functions or demand labor-intensive expert evaluations. Recent progress in Multimodal Large Language Models (MLLMs) reveals their strong visual reasoning capabilities, making them adept candidates as evaluators. In this work, we propose a novel framework, coined MAGIC (Medically Accurate Generation of Images through AI-Expert Collaboration), that synthesizes clinically accurate skin disease images for data augmentation. Our method creatively translates expert-defined criteria into actionable feedback for image synthesis of DMs, significantly improving clinical accuracy while reducing the direct human workload. Experiments demonstrate that our method greatly improves the clinical quality of synthesized skin disease images, with outputs aligning with dermatologist assessments. Additionally, augmenting training data with these synthesized images improves diagnostic accuracy by +9.02% on a challenging 20-condition skin disease classification task, and by +13.89% in the few-shot setting. Beyond image synthesis, MAGIC illustrates a task-centric alignment paradigm: instead of adapting MLLMs to niche medical tasks, it adapts tasks to the evaluative strengths of general-purpose MLLMs by decomposing domain knowledge into attribute-level checklists. This design offers a scalable and reliable path for leveraging foundation models in specialized domains. Our implementation detail and code is available at https://github.com/janet-sw/MAGIC.git. "An image of lupus erythematosus" AI-Expert Collaboration Ground truth: sarcoidosis 5/5 Score 2/5 Score Lupus: 1. swelling or rashes 2. butterfly rash across cheeks 3. scaly or scarred 4. … (a). Clinical Feedback Curation (b). Fine-tune DM with Feedback (c). Train Classifier with Augmented Data
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI UnderstandingYuxiang Wei, Yanteng Zhang, Xi Xiao, Chengxuan Qian 等CVPR 2026 · 被引用 11 次
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization PerspectiveJustin Lee, Zheda Mai, Jinsu Yoo, Chongyu Fan 等ICLR 2026 · 被引用 9 次
- Prime Once, then Reprogram Locally: An Efficient Alternative to Black-Box Service Model AdaptationYunbei Zhang, Chengyi Cai, Feng Liu, Jihun HammCVPR 2026 · 被引用 5 次
它引用的顶会 Paper22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- Expert-guided Clinical Text Augmentation via Query-Based Model CollaborationDongkyu Cho, Miao Zhang, Gregory Lyng, Rumi ChunaraICML 2026
- Aligning Synthetic Medical Images with Clinical Knowledge using Human FeedbackShenghuan Sun, Gregory M. Goldgof, Atul J. Butte, Ahmed M. AlaaNeurIPS 2023 · 被引用 27 次
- MISF: MLLM Guided Iterative Sample Filtering for Data Fault DetectionGuoying Chen, Ruizhuo Zhao, Zhewei Xu, Bo Yang 等AAAI 2026
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical ReasoningPeng Xia, Jinglu Wang, Yibo Peng, Kaide Zeng 等ICLR 2026 · 被引用 47 次
- MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from TextbooksWenqi Zeng, Yuqi Sun, Chenxi Ma, Weimin Tan 等ACM MM 2025 · 被引用 5 次
