Localizing Object-level Shape Variations with Text-to-Image Diffusion Models
Or Patashnik, Daniel Garibi, Idan Azuri, Hadar Averbuch-Elor, Daniel Cohen-Or
Abstract
Text-to-image models give rise to workflows which often begin with an exploration step, where users sift through a large collection of generated images. The global nature of the text-to-image generation process prevents users from narrowing their exploration to a particular object in the image. In this paper, we present a technique to generate a collection of images that depicts variations in the shape of a specific object, enabling an object-level shape exploration process. Creating plausible variations is challenging as it requires control over the shape of the generated object while respecting its semantics. A particular challenge when generating object variations is accurately localizing the manipulation applied over the object’s shape. We introduce a prompt-mixing technique that switches between prompts along the denoising process to attain a variety of shape choices. To localize the image-space operation, we present two techniques that use the self-attention layers in conjunction with the cross-attention layers. Moreover, we show that these localization techniques are general and effective beyond the scope of generating object variations. Extensive results and comparisons demonstrate the effectiveness of our method in generating object variations, and the competence of our localization techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers86
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- FreeNoise: Tuning-Free Longer Video Diffusion via Noise ReschedulingHaonan Qiu, Menghan Xia, Yong Zhang, Yingqing He et al.ICLR 2024 · 171 citations
- Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingKai Wang, Fei Yang, Shiqi Yang, Muhammad Atif Butt et al.NeurIPS 2023 · 108 citations
- Expressive Text-to-Image Generation with Rich TextSongwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin HuangICCV 2023 · 102 citations
- Cross-Image Attention for Zero-Shot Appearance TransferYuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor et al.SIGGRAPH 2024 · 72 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman et al.ICLR 2023 · 361 citations
- Conditional Score Guidance for Text-Driven Image-to-Image TranslationHyunsoo Lee, Minsoo Kang, Bohyung HanNeurIPS 2023 · 23 citations
- Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion ModelsKatarzyna Zaleska, Lukasz Popek, Monika Wysoczanska, Kamil DejaCVPR 2026 · 2 citations
- Diffusion Self-Guidance for Controllable Image GenerationDave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros et al.NeurIPS 2023 · 411 citations
- Zero-shot Image-to-Image TranslationGaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li et al.SIGGRAPH 2023 · 355 citations
