Textualize Visual Prompt for Image Editing via Diffusion Bridge
Pengcheng Xu, Qingnan Fan, Fei Kou, Shuai Qin, Hong Gu, Ruoyu Zhao, Charles Ling, Boyu Wang
摘要
Visual prompt, a pair of before-and-after edited images, can convey indescribable imagery transformations and prosper in image editing. However, current visual prompt methods rely on a pretrained text-guided image-to-image generative model that requires a triplet of text, before, and after images for retraining over a text-to-image model. Such crafting triplets and retraining processes limit the scalability and generalization of editing. In this paper, we present a framework based on any single text-to-image model without reliance on the explicit image-to-image model thus enhancing the generalizability and scalability. Specifically, by leveraging the probability-flow ordinary equation, we construct a diffusion bridge to transfer the distribution between before-and-after images under the text guidance. By optimizing the text via the bridge, the framework adaptively textualizes the editing transformation conveyed by visual prompts into text embeddings without other models. Meanwhile, we introduce differential attention control during text optimization, which disentangles the text embedding from the invariance of the before-and-after images and makes it solely capture the delicate transformation and generalize to edit various images. Experiments on real images validate competitive results on the generalization, contextual coherence, and high fidelity for delicate editing with just one image pair as the visual prompt.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ChordEdit: One-Step Low-Energy Transport for Image EditingLiangsi Lu, Xuhang Chen, Minzhe Guo, Shichu Li 等CVPR 2026 · 被引用 22 次
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative AnalysisKaizhen Zhu, Mokai Pan, Zhechuan Yu, Jingya Wang 等ICML 2026 · 被引用 3 次
- From Discriminative to Generative: A Diffusion-Based Paradigm for Multi-Agent Collaborative PerceptionKexin Gong, Puyi Yao, Guiyang Luo, Quan Yuan 等AAAI 2026
它引用的顶会 Paper41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Flowedit: Inversion-Free Text-Based Editing Using Pre-Trained Flow ModelsVladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas, Tomer MichaeliICCV 2025 · 被引用 30 次
- Training-Free Text-Guided Image Editing with Visual Autoregressive ModelYufei Wang, Lanqing Guo, Zhihao Li, Jiaxing Huang 等ICCV 2025
- An Item Is Worth a Prompt: Versatile Image Editing with Disentangled ControlAosong Feng, Weikang Qiu, Jinbin Bai, Zhen Dong 等AAAI 2025 · 被引用 9 次
- Semantic Editing with Coupled Stochastic Differential EquationsJianxin Zhang, Clay ScottICML 2026
- Video-P2P: Video Editing with Cross-Attention ControlShaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin 等CVPR 2024 · 被引用 99 次
