Is Visual in-Context Learning for Compositional Medical Tasks Within Reach?
Simon Reiß, Zdravko Marinov, Alexander Jaus, Constantin Seibold, M. Saquib Sarfraz, Erik Rodner, Rainer Stiefelhagen
Abstract
In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than individual tasks. Our goal is to solve complex tasks that involve multiple intermediate steps using a single model, allowing users to define entire vision pipelines flexibly at test time. To achieve this, we first examine the properties and limitations of visual in-context learning architectures, with a particular focus on the role of codebooks. We then introduce a novel method for training in-context learners using a synthetic compositional task generation engine. This engine bootstraps task sequences from arbitrary segmentation datasets, enabling the training of visual in-context learners for compositional tasks. Additionally, we investigate different masking-based training objectives to gather insights into how to train models better for solving complex, compositional tasks. Our exploration not only provides important insights especially for multi-modal medical task sequences but also highlights challenges that need to be addressed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
- xLSTM: Extended Long Short-Term MemoryMaximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer et al.NeurIPS 2024 · 703 citations
Related papers
- Towards More Unified In-Context Visual UnderstandingDianmo Sheng, Dongdong Chen, Zhentao Tan, Qiankun Liu et al.CVPR 2024 · 9 citations
- Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context LearningXinshun Wang, Zhongbin Fang, Xia Li, Xiangtai Li et al.CVPR 2024 · 12 citations
- Visual in-Context PromptingFeng Li, Qing Jiang, Hao Zhang, Tianhe Ren et al.CVPR 2024
- Towards Robust Sequential Decomposition for Complex Image EditingZilai Zeng, Mingdeng Cao, Zijie Li, Xiaochen Lian et al.CVPR 2026
- VINCIE: Unlocking In-context Image Editing from VideoLeigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao et al.ICLR 2026 · 18 citations
