Exploring Task-Level Optimal Prompts for Visual In-Context Learning
Yan Zhu, Huan Ma, Changqing Zhang
Abstract
With the development of Vision Foundation Models (VFMs) in recent years, Visual In-Context Learning (VICL) has become a better choice compared to modifying models in most scenarios. Different from retraining or fine-tuning models, VICL does not require modifications to the model's weights and architecture, and only needs a prompt with demonstrations to teach VFM how to solve tasks. Currently, significant computational cost for finding optimal prompts for every test sample hinders the deployment of VICL, as determining which demonstrations to use for constructing the prompt is very costly. In this paper, however, we find a counterintuitive phenomenon that most test samples actually achieve optimal performance under the same prompts, and searching for sample-level prompts only costs much time but results in completely identical prompts actually. Therefore, we propose task-level prompting to reduce the cost of searching for prompts during the inference stage and introduce two time-saving yet effective task-level prompt search strategies accordingly. Extensive experimental results show that our proposed method can identify near-optimal prompts and reach the best VICL performance with a minimal cost that prior work has never achieved.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7987fea-30ab-43e2-9409-bf0dc7cae719Cited by top-tier papers2
- Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context LearningTianci Luo, Haohao Pan, Jinpeng Wang, Niu Lian et al.CVPR 2026
- Efficient and Effective In-context Demonstration Selection with CoresetZihua Wang, Jiarui Wang, Haiyang Xu, Ming Yan et al.AAAI 2026
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Visual Prompting via Image InpaintingAmir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson et al.NeurIPS 2022 · 340 citations
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
- What Makes Good Examples for Visual In-Context Learning?Yuanhan Zhang, Kaiyang Zhou, Ziwei LiuNeurIPS 2023 · 219 citations
Related papers
- Generalizable Object Re-Identification via Visual in-Context PromptingZhizhong Huang, Xiaoming LiuICCV 2025 · 3 citations
- Towards Global Optimal Visual In-Context Learning Prompt SelectionChengming Xu, Chen Liu, Yikai Wang, Yuan Yao et al.NeurIPS 2024 · 19 citations
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen et al.NeurIPS 2024 · 37 citations
- Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningCheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao et al.CVPR 2025
- SINC: Self-Supervised In-Context Learning for Vision-Language TasksYi-Syuan Chen, Yun-Zhu Song, Cheng Yu Yeo, Bei Liu et al.ICCV 2023 · 8 citations
