REMEMBER: Retrieval-based Explainable Multimodal Evidence-guided Modeling for Brain Evaluation and Reasoning in Zero- and Few-shot Neurodegenerative Diagnosis
Duy-Cat Can, Quang-Huy Tang, Huong Ha, Binh T. Nguyen, Oliver Y. Chén
Abstract
Timely and accurate diagnosis of neurodegenerative disorders, such as Alzheimer's disease, is central to disease management. Existing deep learning models require large annotated datasets and often act as ''black boxes''. However, clinical datasets are frequently small or lack labels, limiting the effectiveness of these methods. Here, we introduce REMEMBER - Retrieval-based Explainable Multimodal Evidence-guided Modeling for Brain Evaluation and Reasoning - a machine learning framework that enables zero- and few-shot Alzheimer's diagnosis from brain MRI scans via reference-based reasoning. Specifically, REMEMBER first contrastively trains a vision-text model on expert-annotated reference data, using pseudo-text modalities to encode abnormality types, diagnosis labels, and composite clinical descriptions. At inference time, it retrieves similar, human-validated cases from a curated dataset and integrates their contextual information via an evidence encoder and attention-based inference head. This evidence-guided design allows REMEMBER to mimic clinical decision-making by grounding predictions in retrieved imaging and textual context. It outputs diagnostic predictions with an interpretable report, including reference images and clinical-aligned explanations. Experimental results demonstrate that REMEMBER achieves robust zero- and few-shot performance and offers a powerful and explainable framework to neuroimaging-based diagnosis in the real world, especially under limited data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 717866fb-7e67-4fa7-b288-ca281d9e60a2Builds on4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 907 citations
- GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image RecognitionShih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena YeungICCV 2021 · 516 citations
Related papers
- EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's DiseaseQiuhui Chen, Xuancheng Yao, Zhenglei Zhou, Xinyue Hu et al.CVPR 2026
- SMART: Self-Weighted Multimodal Fusion for Diagnostics of Neurodegenerative DisordersQiuhui Chen, Yi HongACM MM 2024 · 5 citations
- Joint Adaptation of Uni-modal Foundation Models for Multi-modal Alzheimer's Disease DiagnosisWentao Gu, Yuquan Li, Xinyang Jiang, Zilong Wang et al.ICLR 2026
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 1 citation
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical DiagnosisHaolin Li, Tianjie Dai, Zhe Chen, Siyuan Du et al.NeurIPS 2025 · 3 citations
