CLiC: Concept Learning in Context
Mehdi Safaee, Aryan Mikaeili, Or Patashnik, Daniel Cohen-Or, Ali Mahdavi-Amiri
Abstract
This paper addresses the challenge of learning a local visual pattern of an object from one image, and generating images depicting objects with that pattern. Learning a localized concept and placing it on an object in a target image is a nontrivial task, as the objects may have different orientations and shapes. Our approach builds upon recent advancements in visual concept learning. It involves ac-quiring a visual concept (e.g., an ornament) from a source image and subsequently applying it to an object (e.g., a chair) in a target image. Our key idea is to perform in-context concept learning, acquiring the local visual concept within the broader context of the objects they belong to. To localize the concept learning, we employ soft masks that contain both the concept within the mask and the surrounding image area. We demonstrate our approach through object generation within an image, showcasing plausible embedding of in-context learned concepts. We also introduce methods for directing acquired concepts to specific locations within target images, employing cross-attention mechanisms, and establishing correspondences between source and target objects. The effectiveness of our method is demonstrated through quantitative and qualitative experiments, along with comparisons against baseline techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42cbd12d-2114-4e53-accb-010eff564bb2Cited by top-tier papers11
- Zero-shot Image Editing with Reference ImitationXi Chen, Yutong Feng, Mengting Chen, Yiyang Wang et al.NeurIPS 2024 · 80 citations
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven GenerationAbdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter WonkaNeurIPS 2025 · 6 citations
- Magiccolor: Multi-Instance Sketch ColorizationYinhan Zhang, Yue Ma, Bingyuan Wang, Qifeng Chen et al.ICCV 2025 · 3 citations
- PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion ModelsAleksandar Cvejic, Abdelrahman Eldesokey, Peter WonkaSIGGRAPH 2025 · 3 citations
- HOComp: Interaction-Aware Human-Object CompositionDong Liang, Jinyuan Jia, Yuhao Liu, Rynson W. H. LauNeurIPS 2025 · 1 citation
Builds on33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- MaskInversion: Localized Embeddings via Optimization of Explainability MapsWalid Bousselham, Sofian Chaybouti, Christian Rupprecht, Vittorio Ferrari et al.ICLR 2026 · 3 citations
- Customizing Text-to-Image Generation with Inverted InteractionMengmeng Ge, Xu Jia, Takashi Isobe, Xiaomin Li et al.ACM MM 2024 · 3 citations
- Concept-Guided Prompt Learning for Generalization in Vision-Language ModelsYi Zhang, Ce Zhang, Ke Yu, Yushun Tang et al.AAAI 2024 · 37 citations
- Unsupervised Learning of Compositional Energy ConceptsYilun Du, Shuang Li, Yash Sharma, Josh Tenenbaum et al.NeurIPS 2021 · 95 citations
- Latent Expression Generation for Referring Image Segmentation and GroundingSeonghoon Yu, Joonbeom Hong, Joonseok Lee, Jeany SonICCV 2025 · 1 citation
