Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
Xingang Pan, Ayush Tewari, Thomas Leimkühler, Lingjie Liu, Abhimitra Meka, Christian Theobalt
Abstract
Synthesizing visual content that meets users’ needs often requires flexible and precise controllability of the pose, shape, expression, and layout of the generated objects. Existing approaches gain controllability of generative adversarial networks (GANs) via manually annotated training data or a prior 3D model, which often lack flexibility, precision, and generality. In this work, we study a powerful yet much less explored way of controlling GANs, that is, to "drag" any points of the image to precisely reach target points in a user-interactive manner, as shown in Fig.1. To achieve this, we propose DragGAN, which consists of two main components: 1) a feature-based motion supervision that drives the handle point to move towards the target position, and 2) a new point tracking approach that leverages the discriminative generator features to keep localizing the position of the handle points. Through DragGAN, anyone can deform an image with precise control over where pixels go, thus manipulating the pose, shape, expression, and layout of diverse categories such as animals, cars, humans, landscapes, etc. As these manipulations are performed on the learned generative image manifold of a GAN, they tend to produce realistic outputs even for challenging scenarios such as hallucinating occluded content and deforming shapes that consistently follow the object’s rigidity. Both qualitative and quantitative comparisons demonstrate the advantage of DragGAN over prior approaches in the tasks of image manipulation and point tracking. We also showcase the manipulation of real images through GAN inversion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f32c074c-04e9-4937-93c9-fe48e1a39732Cited by top-tier papers107
- DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsChong Mou, Xintao Wang, Jiechong Song, Ying Shan et al.ICLR 2024 · 223 citations
- Understanding the Latent Space of Diffusion Models through the Lens of Riemannian GeometryYong-Hyun Park, Mingi Kwon, Jaewoong Choi, Junghyo Jo et al.NeurIPS 2023 · 163 citations
- PromptMagician: Interactive Prompt Engineering for Text-to-Image CreationYingchaojie Feng, Xingbo Wang, Kamkwai Wong, Sijia Wang et al.IEEE VIS 2023 · 127 citations
- DragDiffusion: Harnessing Diffusion Models for Interactive Point-Based Image EditingYujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan et al.CVPR 2024 · 117 citations
- DirectGPT: A Direct Manipulation Interface to Interact with Large Language ModelsDamien Masson, Sylvain Malacria, Géry Casiez, Daniel VogelCHI 2024 · 104 citations
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
Related papers
- Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive MannerPengxiang Cai, Zhiwei Liu, Guibo Zhu, Yunfang Niu et al.ACM MM 2024 · 3 citations
- EasyDrag: Efficient Point-Based Manipulation on Diffusion ModelsXingzhong Hou, Boxiao Liu, Yi Zhang, Jihao Liu et al.CVPR 2024
- BodyGAN: General-purpose Controllable Neural Human Body GenerationChaojie Yang, Hanhui Li, Shengjie Wu, Shengkai Zhang et al.CVPR 2022 · 8 citations
- GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point DiffusionXiaojian Lin, Hanhui Li, Yuhao Cheng, Yiqiang Yan et al.ICLR 2025
- View Independent Generative Adversarial Network for Novel View SynthesisXiaogang Xu, Ying-Cong Chen, Jiaya JiaICCV 2019 · 43 citations
