What Makes Good Examples for Visual In-Context Learning?
Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu
摘要
Large-scale models trained on broad data have recently become the mainstream architecture in computer vision due to their strong generalization performance. In this paper, the main focus is on an emergent ability in large vision models, known as in-context learning, which allows inference on unseen tasks by conditioning on in-context examples (a.k.a. prompt) without updating the model parameters. This concept has been well-known in natural language processing but has only been studied very recently for large vision models. We for the first time provide a comprehensive investigation on the impact of in-context examples in computer vision, and find that the performance is highly sensitive to the choice of in-context examples. To overcome the problem, we propose a prompt retrieval framework to automate the selection of in-context examples. Specifically, we present (1) an unsupervised prompt retrieval method based on nearest example search using an off-the-shelf model, and (2) a supervised prompt retrieval method, which trains a neural network to choose examples that directly maximize in-context learning performance. The results demonstrate that our methods can bring non-trivial improvements to visual in-context learning in comparison to the commonly-used random selection. The code and models are available at https: //github.com/ZhangYuanhan-AI/ visual_prompt_retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper63
- In-Context Learning Unlocked for Diffusion ModelsZhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen 等NeurIPS 2023 · 被引用 128 次
- OmniSVG: A Unified Scalable Vector Graphics Generation ModelYiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng 等NeurIPS 2025 · 被引用 90 次
- ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image ManipulationYasheng Sun, Yifan Yang, Houwen Peng, Yifei Shen 等NeurIPS 2023 · 被引用 71 次
- Towards In-context Scene UnderstandingIvana Balazevic, David Steiner, Nikhil Parthasarathy, Relja Arandjelovic 等NeurIPS 2023 · 被引用 62 次
- Visual Instruction Inversion: Image Editing via Image PromptingThao Nguyen, Yuheng Li, Utkarsh Ojha, Yong Jae LeeNeurIPS 2023 · 被引用 53 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Fairness-guided Few-shot Prompting for Large Language ModelsHuan Ma, Changqing Zhang, Yatao Bian, Lemao Liu 等NeurIPS 2023 · 被引用 87 次
- How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?Yang Luo, Zangwei Zheng, Zirui Zhu, Yang YouEMNLP 2024 · 被引用 2 次
- Visual in-Context PromptingFeng Li, Qing Jiang, Hao Zhang, Tianhe Ren 等CVPR 2024
- Stable Diffusion Models Are Secretly Good at Visual In-Context LearningTrevine Oorloff, Vishwanath Sindagi, Wele Gedara Chaminda Bandara, Ali Shafahi 等ICCV 2025 · 被引用 11 次
- Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of ExemplarsZhaoxuan Wu, Xiaoqiang Lin, Zhongxiang Dai, Wenyang Hu 等NeurIPS 2024 · 被引用 44 次
