ProConMV: Provenance-Enabled Conceptual Framework for Interpretable Multi-View Diabetic Retinopathy Diagnosis
Xiaoling Luo, Shuo Yang, Qihao Xu, Jiansong Zhang, Zhuoqin Yang, Zhihui Lai, Linlin Shen, Chengliang Liu
Abstract
Existing deep learning models have demonstrated potential in Diabetic retinopathy (DR) diagnosis, but they still suffer from three key challenges: reliance on single-source inputs, opaque and untraceable reasoning processes, and the absence of a mechanism for result verification. Thus, we propose a provenance-enabled concept-based framework for multi-view DR diagnostic (Pro-ConMV), which integrates DR lesion masks, clinical text and multi-view data, utilizing multimodal prompt analysis and visual-text concept interaction to learn the interpretable multi-source input. During the reasoning stage, the proposed framework introduces lesion concepts for causal reasoning chains combining clinical guidelines, and adds doctor intervention for human-machine collaboration. For dynamic fusion decision and verification in multi-view DR diagnosis, we derive via generalization theory that incorporating each view's lesion concept uncertainty and grading uncertainty reduces the generalization error upper bound. Accordingly, we design a dual uncertainty-aware module to enable provenancebased verification, ultimately enabling verifiable analysis of DR diagnostic results. Extensive experiments conducted on two public multi-view DR datasets demonstrate the effectiveness of our method. The code will be released at https: //github.com/SoY0ung/ProConMV.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b53f157-716d-49dc-8e56-2d3b37188c41Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
Related papers
- Vision-Language Models Guided Graph Concept Reasoning for Interpretable Diabetic Retinopathy DiagnosisQihao Xu, Xiaoling Luo, Yuxin Lin, Chengliang Liu et al.AAAI 2026
- Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy GradingXiaoling Luo, Qihao Xu, Huisi Wu, Chengliang Liu et al.AAAI 2025 · 3 citations
- Deep Multi-Task Learning for Diabetic Retinopathy Grading in Fundus ImagesXiaofei Wang, Mai Xu, Jicong Zhang, Lai Jiang et al.AAAI 2021 · 48 citations
- MVCINN: Multi-View Diabetic Retinopathy Detection Using a Deep Cross-Interaction Neural NetworkXiaoling Luo, Chengliang Liu, Waikeung Wong, Jie Wen et al.AAAI 2023 · 14 citations
- Towards Zero-Shot Diabetic Retinopathy Grading: Learning Generalized Knowledge via Prompt-Driven Matching and EmulatingHuan Wang, Haoran Li, Yuxin Lin, Huaming Chen et al.AAAI 2026
