Accelerating Text-to-Image Editing via Cache-Enabled Sparse Diffusion Inference
Zihao Yu, Haoyang Li, Fangcheng Fu, Xupeng Miao, Bin Cui
摘要
Due to the recent success of diffusion models, text-to-image generation is becoming increasingly popular and achieves a wide range of applications. Among them, text-to-image editing, or continuous text-to-image generation, attracts lots of attention and can potentially improve the quality of generated images. It's common to see that users may want to slightly edit the generated image by making minor modifications to their input textual descriptions for several rounds of diffusion inference. However, such an image editing process suffers from the low inference efficiency of many existing diffusion models even using GPU accelerators.
To solve this problem, we introduce Fast Image Semantically Edit (FISEdit), a cached-enabled sparse diffusion model inference engine for efficient text-to-image editing. The key intuition behind our approach is to utilize the semantic mapping between the minor modifications on the input text and the affected regions on the output image. For each text editing step, FISEdit can 1) automatically identify the affected image regions and 2) utilize the cached unchanged regions' feature map to accelerate the inference process. For the former, we measure the differences between cached and ad hoc feature maps given the modified textual description, extract the region with significant differences, and capture the affected region by masks. For the latter, we develop an efficient sparse diffusion inference engine that only computes the feature maps for the affected region while reusing the cached statistics for the rest of the image. Finally, extensive empirical results show that FISEdit can be 3.4 times and 4.4 times faster than existing methods on NVIDIA TITAN RTX and A100 GPUs respectively, and even generates more satisfactory images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Hand1000: Generating Realistic Hands from Text with Only 1, 000 ImagesHaozhuo Zhang, Bin Zhu, Yu Cao, Yanbin HaoAAAI 2025 · 被引用 11 次
- GENTI: GPU-powered Walk-based Subgraph Extraction for Scalable Representation Learning on Dynamic GraphsZihao Yu, Ningyi Liao, Siqiang LuoVLDB 2024 · 被引用 8 次
- Attacks on Approximate Caches in Text-to-Image Diffusion ModelsDesen Sun, Shuncheng Jie, Sihang LiuUSENIX Security 2026 · 被引用 1 次
- Parameterized Complexity of Caching in NetworksRobert Ganian, Fionn Mc Inerney, Dimitra TsigkariAAAI 2025
- Neuro-3D: Towards 3D Visual Decoding from EEG SignalsZhanqiang Guo, Jiamin Wu, Yonghao Song, Jiahui Bu 等CVPR 2025
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 被引用 102 次
- Blended Latent DiffusionOmri Avrahami, Ohad Fried, Dani LischinskiSIGGRAPH 2023 · 被引用 339 次
- Efficient Spatially Sparse Inference for Conditional GANs and Diffusion ModelsMuyang Li, Ji Lin, Chenlin Meng, Stefano Ermon 等NeurIPS 2022 · 被引用 66 次
- Towards Efficient Diffusion-Based Image Editing with Instant Attention MasksSiyu Zou, Jiji Tang, Yiyi Zhou, Jing He 等AAAI 2024 · 被引用 24 次
- Distraction is All You Need: Memory-Efficient Image Immunization against Diffusion-Based Image EditingLing Lo, Cheng Yu Yeo, Hong-Han Shuai, Wen-Huang ChengCVPR 2024 · 被引用 5 次
