Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model
Zheng Gu, Shiyuan Yang, Jing Liao, Jing Huo, Yang Gao
Abstract
Visual In-Context Learning (ICL) has emerged as a promising research area due to its capability to accomplish various tasks with limited example pairs through analogical reasoning. However, training-based visual ICL has limitations in its ability to generalize to unseen tasks and requires the collection of a diverse task dataset. On the other hand, existing methods in the inference-based visual ICL category solely rely on textual prompts, which fail to capture fine-grained contextual information from given examples and can be time-consuming when converting from images to text prompts. To address these challenges, we propose Analogist, a novel inference-based visual ICL approach that exploits both visual and textual prompting techniques using a text-to-image diffusion model pretrained for image inpainting. For visual prompting, we propose a self-attention cloning (SAC) method to guide the fine-grained structural-level analogy between image examples. For textual prompting, we leverage GPT-4V's visual reasoning capability to efficiently generate text prompts and introduce a cross-attention masking (CAM) operation to enhance the accuracy of semantic-level analogy guided by text prompts. Our method is out-of-the-box and does not require fine-tuning or optimization. It is also generic and flexible, enabling a wide range of visual tasks to be performed in an in-context manner. Extensive experiments demonstrate the superiority of our method over existing approaches, both qualitatively and quantitatively. Our project webpage is available at https://analogist2d.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- ImageRAG: Dynamic Image Retrieval for Reference-Guided Image GenerationRotem Shalev-Arkushin, Rinon Gal, Amit Bermano, Ohad FriedICLR 2026 · 25 citations
- PairEdit: Learning Semantic Variations for Exemplar-based Image EditingHaoguang Lu, Jiacheng Chen, Zhenguo Yang, Aurele Tohokantche Gnanha et al.NeurIPS 2025 · 8 citations
- Omni-3DEdit: Generalized Versatile 3D Editing in One-PassLiyi Chen, Pengfei Wang, Guowen Zhang, Zhiyuan Ma et al.CVPR 2026 · 6 citations
- Textualize Visual Prompt for Image Editing via Diffusion BridgePengcheng Xu, Qingnan Fan, Fei Kou, Shuai Qin et al.AAAI 2025 · 4 citations
- Splatent: Splatting Diffusion Latents for Novel View SynthesisOr Hirschorn, Omer Sela, Inbar Huberman-Spiegelglas, Netalee Efrat Sela et al.CVPR 2026 · 2 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
Related papers
- Stable Diffusion Models Are Secretly Good at Visual In-Context LearningTrevine Oorloff, Vishwanath Sindagi, Wele Gedara Chaminda Bandara, Ali Shafahi et al.ICCV 2025 · 11 citations
- In-Context Learning Unlocked for Diffusion ModelsZhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen et al.NeurIPS 2023 · 128 citations
- VisualCloze: A Universal Image Generation Framework via Visual in-Context LearningZhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo et al.ICCV 2025 · 2 citations
- ConText: Driving In-context Learning for Text Removal and SegmentationFei Zhang, Pei Zhang, Baosong Yang, Fei Huang et al.ICML 2025
- Towards Global Optimal Visual In-Context Learning Prompt SelectionChengming Xu, Chen Liu, Yikai Wang, Yuan Yao et al.NeurIPS 2024 · 19 citations
