DragNeXt: Rethinking Drag-Based Image Editing
Yuan Zhou, Junbao Zhou, Qingshan Xu, Kesen Zhao, Yuxuan Wang, Hao Fei, Richang Hong, Hanwang Zhang
摘要
Drag-Based Image Editing (DBIE), which allows users to manipulate images by directly dragging objects within them, has recently attracted much attention from the community. However, it faces two key challenges: (i) point-based drag is often highly ambiguous and difficult to align with user intentions; (ii) current DBIE methods primarily rely on alternating between motion supervision and point tracking, which is not only cumbersome but also fails to produce high-quality results. These limitations motivate us to explore DBIE from a new perspective---unifying it as a Latent Region Optimization (LRO) problem that aims to use region-level geometric transformations to optimize latent code to realize drag manipulation. Thus, by specifying the areas and types of geometric transformations, we can effectively address the ambiguity issue. We also propose a simple yet effective editing framework, dubbed DragNeXt. It solves LRO through Progressive Backward Self-Intervention (PBSI), simplifying the overall procedure of the alternating workflow while further enhancing quality by fully leveraging region-level structure information and progressive guidance from intermediate drag states. We validate DragNeXt on our NextBench, and extensive experiments demonstrate that our proposed method can significantly outperform existing approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point DiffusionXiaojian Lin, Hanhui Li, Yuhao Cheng, Yiqiang Yan 等ICLR 2025
- FastDrag: Manipulate Anything in One StepXuanjia Zhao, Jian Guan, Congyi Fan, Dongli Xu 等NeurIPS 2024 · 被引用 27 次
- Dragging with Geometry: From Pixels to Geometry-Guided Image EditingXinyu Pu, Hongsong Wang, Jie Gui, Pan ZhouICLR 2026 · 被引用 5 次
- CLIPDrag: Combining Text-based and Drag-based Instructions for Image EditingZiqi Jiang, Zhen Wang, Long ChenICLR 2025
- FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow FieldsGwanhyeong Koo, Sunjae Yoon, Younghwan Lee, Ji Woo Hong 等ICML 2025
