Towards Zero-Shot Diabetic Retinopathy Grading: Learning Generalized Knowledge via Prompt-Driven Matching and Emulating
Huan Wang, Haoran Li, Yuxin Lin, Huaming Chen, Jun Yan, Lijuan Wang, Jiahua Shi, Qihao Xu, Yongting Hu, Yong Xu, Jun Shen
Abstract
As one of the primary causes of visual impairment, Diabetic Retinopathy (DR) requires accurate and robust grading to facilitate timely diagnosis and intervention. Different from conventional DR grading methods that utilize single-view images, recent clinical studies have revealed that multi-view fundus images can significantly enhance DR grading performance by expanding the field of view (FOV). However, there is a long-tailed distribution problem in fundus image analysis, i.e., a high prevalence of mild DR grades and a low prevalence of rare ones (e.g., cases of high severity), which presents a significant challenge to developing a unified model capable of detecting rare or unseen DR grades not encountered during training. In this paper, we propose ProME-DR, a Prompt-driven zero-shot DR grading framework, which leverages prompt Matching and Emulating to recognize the unseen DR categories and views beyond the training set. ProME-DR disentangles the training process into two stages to learn generalized knowledge for novel DR disease grading. Initially, ProME-DR leverages two sets of prompt units to capture semantic and inter-view consistency knowledge via a split-and-mask manner, gathering instance-level DR visual clues. Subsequently, it constructs a concept-aware emulator to generate context prompt units, linking extensible knowledge learned from the previously seen DR attributes for zero-shot DR grading. Extensive experiments conducted on eight datasets and various scenarios confirm the superiority of ProME-DR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy GradingXiaoling Luo, Qihao Xu, Huisi Wu, Chengliang Liu et al.AAAI 2025 · 3 citations
- MVCINN: Multi-View Diabetic Retinopathy Detection Using a Deep Cross-Interaction Neural NetworkXiaoling Luo, Chengliang Liu, Waikeung Wong, Jie Wen et al.AAAI 2023 · 14 citations
- Incomplete Multi-view Diabetic Retinopathy Grading via Self-Supervised Inter- and Intra-View RestorationZhihao Wu, Yuxin Lin, Jie Wen, Wuzhen Shi et al.AAAI 2026
- ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangling and Self-PromptingYankai Jiang, Zhongzhen Huang, Rongzhao Zhang, Xiaofan Zhang et al.CVPR 2024
- Deep Multi-Task Learning for Diabetic Retinopathy Grading in Fundus ImagesXiaofei Wang, Mai Xu, Jicong Zhang, Lai Jiang et al.AAAI 2021 · 48 citations
