HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Education
Qian Wu, Zheyao Gao, Longfei Gou, Yifan Hou, Ann Sin Nga Lau, Qi Dou
Abstract
The evolution of text-to-image (T2I) generation techniques has introduced new capabilities for information visualization, with the potential to advance knowledge democratization and education. In this paper, we investigate how T2I models can be adapted to generate educational health knowledge contents, exploring their potential to make healthcare information more visually accessible and engaging. We explore methods to harness recent T2I models for generating health knowledge flashcards-visual educational aids that present healthcare information through appealing and concise imagery. To support this goal, we curated a diverse, highquality healthcare knowledge flashcard dataset containing 2,034 samples sourced from credible medical resources. We further validate the effectiveness of fine-tuning open-source models with our dataset, demonstrating their promise as specialized health flashcard generators. Our code and dataset are available at: https://github.com/med-air/HealthCards .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- PlankAssembly: Robust 3D Reconstruction from Three Orthographic Views with Learnt Shape ProgramsWentao Hu, Jia Zheng, Zixin Zhang, Xiaojun Yuan et al.ICCV 2023 · 12 citations
Related papers
- Enhancing Textbooks with Visuals from the Web for Improved LearningJanvijay Singh, Vilém Zouhar, Mrinmaya SachanEMNLP 2023 · 2 citations
- Let the Chart Spark: Embedding Semantic Context into Chart with Text-to-Image Generative ModelShishi Xiao, Suizi Huang, Yue Lin, Yilin Ye et al.IEEE VIS 2023 · 44 citations
- MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation modelsMohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi SoltanolkotabiICLR 2025
- Scaling Down Text Encoders of Text-to-Image Diffusion ModelsLifu Wang, Daqing Liu, Xinchen Liu, Xiaodong HeCVPR 2025
- T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive ConceptsZiwei Huang, Wanggui He, Quanyu Long, Yandi Wang et al.ACL 2025
