GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion
Xiaojian Lin, Hanhui Li, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang
摘要
Recent interactive point-based image manipulation methods have gained considerable attention for being user-friendly. However, these methods still face two types of ambiguity issues that can lead to unsatisfactory outcomes, namely, intention ambiguity which misinterprets the purposes of users, and content ambiguity where target image areas are distorted by distracting elements. To address these issues and achieve general-purpose manipulations, we propose a novel taskaware, training-free framework called GDrag. Specifically, GDrag defines a taxonomy of atomic manipulations, which can be parameterized and combined unitedly to represent complex manipulations, thereby reducing intention ambiguity. Furthermore, GDrag introduces two strategies to mitigate content ambiguity, including an anti-ambiguity dense trajectory calculation method (ADT) and a selfadaptive motion supervision method (SMS). Given an atomic manipulation, ADT converts the sparse user-defined handle points into a dense point set by selecting their semantic and geometric neighbors, and calculates the trajectory of the point set. Unlike previous motion supervision methods relying on a single global scale for low-rank adaption, SMS jointly optimizes point-wise adaption scales and latent feature biases. These two methods allow us to model fine-grained target contexts and generate precise trajectories. As a result, GDrag consistently produces precise and appealing results in different editing tasks. Extensive experiments on the challenging DragBench dataset demonstrate that GDrag outperforms state-of-the-art methods significantly. The code of GDrag is available at https://github.com/DaDaY-coder/GDrag.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DragFlow: Unleashing DiT Priors with Region-Based Supervision for Drag EditingZihan Zhou, Shilin Lu, Shuli Leng, Shaocong Zhang 等ICLR 2026 · 被引用 33 次
- FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language ModelJun Zhou, Jiahao Li, Zunnan Xu, Hanhui Li 等CVPR 2025
- DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion ModelSiwei Xia, Li Sun, Tiantian Sun, Qingli LiICML 2025
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- DragNeXt: Rethinking Drag-Based Image EditingYuan Zhou, Junbao Zhou, Qingshan Xu, Kesen Zhao 等AAAI 2026 · 被引用 7 次
- CLIPDrag: Combining Text-based and Drag-based Instructions for Image EditingZiqi Jiang, Zhen Wang, Long ChenICLR 2025
- Drag Your GAN: Interactive Point-based Manipulation on the Generative Image ManifoldXingang Pan, Ayush Tewari, Thomas Leimkühler, Lingjie Liu 等SIGGRAPH 2023 · 被引用 206 次
- FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow FieldsGwanhyeong Koo, Sunjae Yoon, Younghwan Lee, Ji Woo Hong 等ICML 2025
- GoodDrag: Towards Good Practices for Drag Editing with Diffusion ModelsZewei Zhang, Huan Liu, Jun Chen, Xiangyu XuICLR 2025
