Stable Diffusion Models Are Secretly Good at Visual In-Context Learning
Trevine Oorloff, Vishwanath Sindagi, Wele Gedara Chaminda Bandara, Ali Shafahi, Amin Ghiasi, Charan Prakash, Reza Ardekani
Abstract
Large language models (LLM) in natural language processing (NLP) have demonstrated great potential for in-context learning (ICL) - the ability to leverage a few sets of example prompts to adapt to various tasks without having to explicitly update the model weights. ICL has recently been explored for computer vision tasks with promising early outcomes. These approaches involve specialized training and/or additional data that complicate the process and limit its generalizability. In this work, we show that off-the-shelf Stable Diffusion models can be repurposed for visual incontext learning (V-ICL). Specifically, we formulate an inplace attention re-computation within the self-attention layers of the Stable Diffusion architecture that explicitly incorporates context between the query and example prompts. Without any additional fine-tuning, we show that this repurposed Stable Diffusion model is able to adapt to six different tasks: foreground segmentation, single object detection, semantic segmentation, keypoint detection, edge detection, and colorization. For example, the proposed approach improves the mean intersection over union (mIoU) for the foreground segmentation task on Pascal-5i dataset by 8.9% and 3.2% over recent methods such as Visual Prompting and IMProv, respectively. Additionally, we show that the proposed method is able to effectively leverage multiple prompts through ensembling to infer the task better and further improve the performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d33796b0-aa27-4af7-901f-e22df87e763eCited by top-tier papers4
- PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and AlignmentTianci Luo, Jinpeng Wang, Shiyu Qin, Niu Lian et al.ICLR 2026 · 3 citations
- Splatent: Splatting Diffusion Latents for Novel View SynthesisOr Hirschorn, Omer Sela, Inbar Huberman-Spiegelglas, Netalee Efrat Sela et al.CVPR 2026 · 2 citations
- Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context LearningTianci Luo, Haohao Pan, Jinpeng Wang, Niu Lian et al.CVPR 2026
- The Latent Color Subspace: Emergent Order in High-Dimensional ChaosMateusz Pach, Jessica Bader, Quentin Bouniot, Serge Belongie et al.ICML 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan et al.ICCV 2023 · 770 citations
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang et al.NeurIPS 2024 · 728 citations
Related papers
- In-Context Learning Unlocked for Diffusion ModelsZhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen et al.NeurIPS 2023 · 128 citations
- Visual in-Context PromptingFeng Li, Qing Jiang, Hao Zhang, Tianhe Ren et al.CVPR 2024
- SLiMe: Segment Like MeAliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi, Ali Mahdavi-Amiri et al.ICLR 2024 · 47 citations
- What Makes Good Examples for Visual In-Context Learning?Yuanhan Zhang, Kaiyang Zhou, Ziwei LiuNeurIPS 2023 · 219 citations
- Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingKai Wang, Fei Yang, Shiqi Yang, Muhammad Atif Butt et al.NeurIPS 2023 · 108 citations
