FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields
Gwanhyeong Koo, Sunjae Yoon, Younghwan Lee, Ji Woo Hong, Chang D. Yoo
Abstract
Drag-based editing allows precise object manipulation through point-based control, offering user convenience. However, current methods often suffer from a geometric inconsistency problem by focusing exclusively on matching user-defined points, neglecting the broader geometry and leading to artifacts or unstable edits. We propose FlowDrag, which leverages geometric information for more accurate and coherent transformations. Our approach constructs a 3D mesh from the image, using an energy function to guide mesh deformation based on user-defined drag points. The resulting mesh displacements are projected into 2D and incorporated into a UNet denoising process, enabling precise handle-to-target point alignment while preserving structural integrity. Additionally, existing drag-editing benchmarks provide no ground truth, making it difficult to assess how accurately the edits match the intended transformations. To address this, we present VFD (VidFrameDrag) benchmark dataset, which provides ground-truth frames using consecutive shots in a video dataset. FlowDrag outperforms existing drag-based editing methods on both VFD Bench and DragBench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Dragging with Geometry: From Pixels to Geometry-Guided Image EditingXinyu Pu, Hongsong Wang, Jie Gui, Pan ZhouICLR 2026 · 5 citations
- Generative Blocks World: Moving Things Around in PicturesVaibhav Vavilala, Seemandhar Jain, Rahul Vasanth, David Forsyth et al.ICLR 2026 · 4 citations
- A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion TransformersTrung X. Pham, Kang Zhang, Ji Woo Hong, Chang Dong YooICLR 2026 · 2 citations
- Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement LearningMinh-Tung Luu, Hwanhee Kim, Younghwan Lee, Chang D. YooICML 2026 · 1 citation
- GADA: Geometry-Aware Deformable Aggregation for Image-Based Gaussian SplattingSiwoo Lim, Sunjae Yoon, Gwanhyeong Koo, Chang D. YooICML 2026
Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- DragFlow: Unleashing DiT Priors with Region-Based Supervision for Drag EditingZihan Zhou, Shilin Lu, Shuli Leng, Shaocong Zhang et al.ICLR 2026 · 33 citations
- GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image GenerationPhillip Mueller, Talip Uenlue, Sebastian Schmidt, Marcel Kollovieh et al.ICCV 2025 · 2 citations
- 3DGS-Drag: Dragging Gaussians for Intuitive Point-Based 3D EditingJiahua Dong, Yu-Xiong WangICLR 2025
- GoodDrag: Towards Good Practices for Drag Editing with Diffusion ModelsZewei Zhang, Huan Liu, Jun Chen, Xiangyu XuICLR 2025
- LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit CorrespondenceZixin Yin, Xili Dai, Duomin Wang, Xianfang Zeng et al.ICLR 2026 · 4 citations
