Evidential Interactive Learning for Medical Image Captioning
Ervine Zheng, Qi Yu
Abstract
Medical image captioning alleviates the burden of physicians and possibly reduces medical errors by automatically generating text descriptions to describe image contents and convey findings. It is more challenging than conventional image captioning due to the complexity of medical images and the difficulty of aligning image regions with medical terms. In this paper, we propose an evidential interactive learning framework that leverages evidence-based uncertainty estimation and interactive machine learning to improve image captioning with limited labeled data. The interactive learning process involves three stages: keyword prediction, caption generation, and model updates. First, the model predicts a list of keywords with evidence-based uncertainty estimation and selects the most informative keywords to seek user feedback. Second, user-approved keywords are used as model input to guide the model to generate satisfactory captions. Third, the model is updated based on user-approved keywords and captions, where evidence-based uncertainty is used to allocate different weights to different data instances. Experiments on two medical image datasets illustrate that the proposed framework can effectively learn from human feedback and improve performance in the future.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- A Continual Learning Framework for Uncertainty-Aware Interactive Image SegmentationErvine Zheng, Qi Yu, Rui Li, Pengcheng Shi et al.AAAI 2021 · 27 citations
- Hierarchical Multi-Source Uncertainty Aggregation for Interactive Video CaptioningErvine Zheng, Qi YuAAAI 2025
- SPA: Efficient User-Preference Alignment against Uncertainty in Medical Image SegmentationJiayuan Zhu, Junde Wu, Cheng Ouyang, Konstantinos Kamnitsas et al.ICCV 2025 · 1 citation
- Multi-Mode Interactive Image SegmentationZheng Lin, Zhao Zhang, Linghao Han, Shao-Ping LuACM MM 2022 · 8 citations
- Evidential Prototype Learning for Semi-supervised Medical Image SegmentationYuanpeng He, Lijian Li, Tianxiang Zhan, Chi-Man Pun et al.KDD 2025 · 1 citation
