"Why Is There a Tumor?": Tell Me the Reason, Show Me the Evidence
Mengmeng Ma, Tang Li, Yunxiang Peng, Lu Lin, Volkan Beylergil, Binsheng Zhao, Oguz Akin, Xi Peng
Abstract
Medical AI models excel at tumor detection and segmentation. However, their latent representations often lack explicit ties to clinical semantics, producing outputs less trusted in clinical practice. Most of the existing models generate either segmentation masks/labels (localizing where without why) or textual justifications (explaining why without where), failing to ground clinical concepts in spatially localized evidence. To bridge this gap, we propose to develop models that can justify the segmentation or detection using clinically relevant terms and point to visual evidence. We address two core challenges: First, we curate a rationale dataset to tackle the lack of paired images, annotations, and textual rationales for training. The dataset includes 180K image-mask-rationale triples with quality evaluated by expert radiologists. Second, we design rationale-informed optimization that disentangles and localizes finegrained clinical concepts in a self-supervised manner without requiring pixel-level concept annotations. Experiments across medical benchmarks show our model demonstrates superior performance in segmentation, detection, and beyond.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- Are Multimodal Transformers Robust to Missing Modality?Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine et al.CVPR 2022 · 153 citations
Related papers
- Beyond Accuracy: Ensuring Correct Predictions With Correct RationalesTang Li, Mengmeng Ma, Xi PengNeurIPS 2024 · 6 citations
- Report-Concept Textual-Prompt Learning for Enhancing X-ray DiagnosisXiongjun Zhao, Zhengyu Liu, Fen Liu, Guanting Li et al.ACM MM 2024 · 3 citations
- TumorChain: Interleaved Multimodal Chain-of-Thought Reasoning for Traceable Clinical Tumor AnalysisSijing Li, Zhongwei Qiu, Jiang Liu, Wenqiao Zhang et al.ICLR 2026 · 2 citations
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing et al.AAAI 2026 · 4 citations
- Medical Vision-Language Pretraining with LLM-Guided Temporal SupervisionLiang Bai, Zhi Wang, Huimin Yan, Xian YangAAAI 2026
