Retriever Encoder Selection Matters for In-Context Learning-based Medical Segmentation
Fan Wang, Zhongyi Han, Yongshun Gong, Yilong Yin
Abstract
In-context learning-based medical segmentation (ICLM) enables foundation models to generalize to unseen cases without retraining. To enhance performance on test queries, existing methods typically follow a two-stage process: (1) using a retrieval encoder (RE) to map both queries and training samples into a shared feature space, and (2) retrieving and utilizing the top-k most similar training samples. While current methods fix the RE and focus on optimizing stage (2), we show that the choice of RE in stage (1) alone can account for over 70% of the performance variation, highlighting RE selection as a critical yet often overlooked factor in ICLM. In this paper, we conduct an analysis of the RE selection and make two main findings: (1) dynamically selecting the RE for each query outperforms selecting a fixed RE for the entire task; and (2) feature-space heuristics (e.g., intra-class compactness and inter-class separability) fail to predict RE quality. To this end, we propose the instance-adaptive retrieval encoder selection (IRES) method that can select the optimal RE for each query based on output predictions. IRES is based on the intuition that a good RE retrieves relevant demonstrations, helping the ICL model generate more accurate and stable segmentation masks. Thus, we introduce the shape stability score (S 3 ), which evaluates the morphological stability of predicted masks under iterative erosion. Experiments show S 3 correlates strongly with true RE quality (Pearson > 0.8), serving as a reliable selection proxy. To reduce S 3 's per-query cost, we propose parallel prediction with reciprocal neighbor reuse (P2R), which accelerates inference by parallelizing encoding and reusing encoder selections across reciprocal neighbors, avoiding redundant computation. Built on S 3 and P2R, IRES improves ICLM performance across FUNDUS, Brain MRI, and Chest X-ray datasets, with up to 10.6% gain on fundus segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 954b8445-8ce4-4182-a8ae-53ce528e4118Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui et al.ICLR 2022 · 565 citations
- Exploiting the Intrinsic Neighborhood Structure for Source-free Domain AdaptationShiqi Yang, Yaxing Wang, Joost van de Weijer, Luis Herranz et al.NeurIPS 2021 · 371 citations
Related papers
- Show and Segment: Universal Medical Image Segmentation via In-Context LearningYunhe Gao, Di Liu, Zhuowei Li, Yunsheng Li et al.CVPR 2025
- DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical ImagingYuxi Ma, Sujie Liu, Jing Yang, Jiacheng Wang et al.CVPR 2026
- Unified Medical Lesion Segmentation via Self-referring IndicatorShijie Chang, Xiaoqi Zhao, Lihe Zhang, Tiancheng WangCVPR 2025
- CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image AnalysisDi Zhang, Zhangpeng Gong, Xiaobo Pang, Jiashuai Liu et al.CVPR 2026 · 11 citations
- Medverse: A Universal Model for Full-Resolution 3D Medical Image Segmentation, Transformation and EnhancementJiesi Hu, Jianfeng Cao, Yanwu Yang, Chenfei Ye et al.AAAI 2026 · 2 citations
