CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
Ziqi Jiang, Zhen Wang, Long Chen
摘要
Precise and flexible image editing remains a fundamental challenge in computer vision. Based on the modified areas, most editing methods can be divided into two main types: global editing and local editing. In this paper, we choose the two most common editing approaches (text-based editing and drag-based editing) and analyze their drawbacks. Specifically, text-based methods often fail to describe the desired modifications precisely, while drag-based methods suffer from ambiguity. To address these issues, we proposed CLIPDrag, a novel image editing method that is the first to combine text and drag signals for precise and ambiguity-free manipulations on diffusion models. To fully leverage these two signals, we treat text signals as global guidance and drag points as local information. Then we introduce a novel global-local motion supervision method to integrate text signals into existing drag-based methods by adapting a pre-trained language-vision model like CLIP. Furthermore, we also address the problem of slow convergence in CLIPDrag by presenting a fast point-tracking method that enforces drag points moving toward correct directions. Extensive experiments demonstrate that CLIPDrag outperforms existing single drag-based methods or text-based methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DragFlow: Unleashing DiT Priors with Region-Based Supervision for Drag EditingZihan Zhou, Shilin Lu, Shuli Leng, Shaocong Zhang 等ICLR 2026 · 被引用 33 次
- DragNeXt: Rethinking Drag-Based Image EditingYuan Zhou, Junbao Zhou, Qingshan Xu, Kesen Zhao 等AAAI 2026 · 被引用 7 次
- Neural-Driven Image EditingPengfei Zhou, Jie Xia, Xiaopeng Peng, Wangbo Zhao 等NeurIPS 2025 · 被引用 5 次
- LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit CorrespondenceZixin Yin, Xili Dai, Duomin Wang, Xianfang Zeng 等ICLR 2026 · 被引用 4 次
- SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video GenerationKien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 被引用 670 次
- DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsChong Mou, Xintao Wang, Jiechong Song, Ying Shan 等ICLR 2024 · 被引用 223 次
- DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image EditingChong Mou, Xintao Wang, Jiechong Song, Ying Shan 等CVPR 2024 · 被引用 36 次
- PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion ModelsAleksandar Cvejic, Abdelrahman Eldesokey, Peter WonkaSIGGRAPH 2025 · 被引用 3 次
- Blended Latent DiffusionOmri Avrahami, Ohad Fried, Dani LischinskiSIGGRAPH 2023 · 被引用 339 次
