Multi-Modal Multi-Instance Learning for Retinal Disease Recognition
Xirong Li, Yang Zhou, Jie Wang, Hailan Lin, Jianchun Zhao, Dayong Ding, Weihong Yu, Youxin Chen
Abstract
This paper attacks an emerging challenge of multi-modal retinal disease recognition. Given a multi-modal case consisting of a color fundus photo (CFP) and an array of OCT B-scan images acquired during an eye examination, we aim to build a deep neural network that recognizes multiple vision-threatening diseases for the given case. As the diagnostic efficacy of CFP and OCT is disease-dependent, the network's ability of being both selective and interpretable is important. Moreover, as both data acquisition and manual labeling are extremely expensive in the medical domain, the network has to be relatively lightweight for learning from a limited set of labeled multi-modal samples. Prior art on retinal disease recognition focuses either on a single disease or on a single modality, leaving multi-modal fusion largely underexplored. We propose in this paper Multi-Modal Multi-Instance Learning (MM-MIL) for selectively fusing CFP and OCT modalities. Its lightweight architecture (as compared to current multi-head attention modules) makes it suited for learning from relatively small-sized datasets. For an effective use of MM-MIL, we propose to generate a pseudo sequence of CFPs by over sampling a given CFP. The benefits of this tactic include well balancing instances across modalities, increasing the resolution of the CFP input, and finding out regions of the CFP most relevant with respect to the final diagnosis. Extensive experiments on a real-world dataset consisting of 1,206 multi-modal cases from 1,193 eyes of 836 subjects demonstrate the viability of the proposed model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 279ab74d-6ff2-4b97-bd3c-20b59156c965Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng et al.ACM MM 2020 · 93 citations
- Relevance-Based Compression of Cataract Surgery Videos Using Convolutional Neural NetworksNegin Ghamsarian, Hadi Amirpour, Christian Timmerer, Mario Taschwer et al.ACM MM 2020 · 20 citations
- Transformer Interpretability Beyond Attention VisualizationHila Chefer, Shir Gur, Lior WolfCVPR 2021
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang et al.CVPR 2020
Related papers
- Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease DiagnosisYuxin Lin, Haoran Li, Haoyu Cao, Yongting Hu et al.AAAI 2026
- DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationHongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao et al.CVPR 2022 · 402 citations
- SI-MIL: Taming Deep MIL for Self-Interpretability in Gigapixel HistopathologySaarthak Kapse, Pushpak Pati, Srijan Das, Jingwei Zhang et al.CVPR 2024
- Bayes-MIL: A New Probabilistic Perspective on Attention-based Multiple Instance Learning for Whole Slide ImagesYufei Cui, Ziquan Liu, Xiangyu Liu, Xue Liu et al.ICLR 2023
- LNPL-MIL: Learning from Noisy Pseudo Labels for Promoting Multiple Instance Learning in Whole Slide ImageZhuchen Shao, Yifeng Wang, Yang Chen, Hao Bian et al.ICCV 2023 · 27 citations
