Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model
Zheng Gu, Shiyuan Yang, Jing Liao, Jing Huo, Yang Gao
摘要
Visual In-Context Learning (ICL) has emerged as a promising research area due to its capability to accomplish various tasks with limited example pairs through analogical reasoning. However, training-based visual ICL has limitations in its ability to generalize to unseen tasks and requires the collection of a diverse task dataset. On the other hand, existing methods in the inference-based visual ICL category solely rely on textual prompts, which fail to capture fine-grained contextual information from given examples and can be time-consuming when converting from images to text prompts. To address these challenges, we propose Analogist, a novel inference-based visual ICL approach that exploits both visual and textual prompting techniques using a text-to-image diffusion model pretrained for image inpainting. For visual prompting, we propose a self-attention cloning (SAC) method to guide the fine-grained structural-level analogy between image examples. For textual prompting, we leverage GPT-4V's visual reasoning capability to efficiently generate text prompts and introduce a cross-attention masking (CAM) operation to enhance the accuracy of semantic-level analogy guided by text prompts. Our method is out-of-the-box and does not require fine-tuning or optimization. It is also generic and flexible, enabling a wide range of visual tasks to be performed in an in-context manner. Extensive experiments demonstrate the superiority of our method over existing approaches, both qualitatively and quantitatively. Our project webpage is available at https://analogist2d.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- ImageRAG: Dynamic Image Retrieval for Reference-Guided Image GenerationRotem Shalev-Arkushin, Rinon Gal, Amit Bermano, Ohad FriedICLR 2026 · 被引用 25 次
- PairEdit: Learning Semantic Variations for Exemplar-based Image EditingHaoguang Lu, Jiacheng Chen, Zhenguo Yang, Aurele Tohokantche Gnanha 等NeurIPS 2025 · 被引用 8 次
- Omni-3DEdit: Generalized Versatile 3D Editing in One-PassLiyi Chen, Pengfei Wang, Guowen Zhang, Zhiyuan Ma 等CVPR 2026 · 被引用 6 次
- Textualize Visual Prompt for Image Editing via Diffusion BridgePengcheng Xu, Qingnan Fan, Fei Kou, Shuai Qin 等AAAI 2025 · 被引用 4 次
- Splatent: Splatting Diffusion Latents for Novel View SynthesisOr Hirschorn, Omer Sela, Inbar Huberman-Spiegelglas, Netalee Efrat Sela 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- Stable Diffusion Models Are Secretly Good at Visual In-Context LearningTrevine Oorloff, Vishwanath Sindagi, Wele Gedara Chaminda Bandara, Ali Shafahi 等ICCV 2025 · 被引用 11 次
- In-Context Learning Unlocked for Diffusion ModelsZhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen 等NeurIPS 2023 · 被引用 128 次
- VisualCloze: A Universal Image Generation Framework via Visual in-Context LearningZhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo 等ICCV 2025 · 被引用 2 次
- ConText: Driving In-context Learning for Text Removal and SegmentationFei Zhang, Pei Zhang, Baosong Yang, Fei Huang 等ICML 2025
- Towards Global Optimal Visual In-Context Learning Prompt SelectionChengming Xu, Chen Liu, Yikai Wang, Yuan Yao 等NeurIPS 2024 · 被引用 19 次
