Cross-Image Attention for Zero-Shot Appearance Transfer
Yuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor, Daniel Cohen-Or
Abstract
Recent advancements in text-to-image generative models have demonstrated a remarkable ability to capture a deep semantic understanding of images. In this work, we leverage this semantic knowledge to transfer the visual appearance between objects that share similar semantics but may differ significantly in shape. To achieve this, we build upon the self-attention layers of these generative models and introduce a cross-image attention mechanism that implicitly establishes semantic correspondences across images. Specifically, given a pair of images — one depicting the target structure and the other specifying the desired appearance — our cross-image attention combines the queries corresponding to the structure image with the keys and values of the appearance image. This operation, when applied during the denoising process, leverages the established semantic correspondences to generate an image combining the desired structure and appearance. In addition, to improve the output image quality, we harness three mechanisms that either manipulate the noisy latent codes or the model’s internal representations throughout the denoising process. Importantly, our approach is zero-shot, requiring no optimization or training. Experiments show that our method is effective across a wide range of object categories and is robust to variations in shape, size, and viewpoint between the two input images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d45e6553-aa87-4752-aa01-0cde2bc50a9bCited by top-tier papers59
- Training-Free Consistent Text-to-Image GenerationYoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten et al.SIGGRAPH 2024 · 57 citations
- Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationJiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang et al.NeurIPS 2025 · 12 citations
- Attention (as Discrete-Time Markov) ChainsYotam Erel, Olaf Dünkel, Rishabh Dabral, Vladislav Golyanik et al.NeurIPS 2025 · 12 citations
- Stable Diffusion Models Are Secretly Good at Visual In-Context LearningTrevine Oorloff, Vishwanath Sindagi, Wele Gedara Chaminda Bandara, Ali Shafahi et al.ICCV 2025 · 11 citations
- Edicho: Consistent Image Editing in the WildQingyan Bai, Hao Ouyang, Yinghao Xu, Qiuyu Wang et al.ICCV 2025 · 10 citations
Builds on45
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Splicing ViT Features for Semantic Appearance TransferNarek Tumanyan, Omer Bar-Tal, Shai Bagon, Tali DekelCVPR 2022 · 128 citations
- Localizing Object-level Shape Variations with Text-to-Image Diffusion ModelsOr Patashnik, Daniel Garibi, Idan Azuri, Hadar Averbuch-Elor et al.ICCV 2023 · 158 citations
- Unsupervised Semantic Correspondence Using Stable DiffusionEric Hedlin, Gopal Sharma, Shweta Mahajan, Hossam Isack et al.NeurIPS 2023 · 152 citations
- UniversalBooth: Model-Agnostic Personalized Text-To-Image GenerationSonghua Liu, Ruonan Yu, Xinchao WangICCV 2025 · 2 citations
- MotionShot: Adaptive Motion Transfer Across Arbitrary Objects for Text-to-Video GenerationYanchen Liu, Yanan Sun, Zhening Xing, Junyao Gao et al.ICCV 2025 · 5 citations
