FSI-Edit: Frequency and Stochasticity Injection for Flexible Diffusion-Based Image Editing
Kaixiang Yang, Xin Li, Yuxi Li, Qiang Li, Zhiwei Wang
摘要
Latent Diffusion-based Text-to-Image (T2I) is a free image editing tool that typically reverses an image into noise, reconstructs it using its original text prompt, and then generates an edited version under a new target prompt. To preserve unaltered image content, features from the reconstruction are directly injected to replace selected features in the generation. However, this direct replacement often leads to feature incompatibility, compromising editing fidelity and limiting creative flexibility, particularly for non-rigid edits ( e.g. , structural or pose changes). In this paper, we aim to address these limitations by proposing FSI-Edit , a novel framework using frequency-and stochasticity-based feature injection for flexible image editing. First, FSI-Edit enhances feature consistency by injecting high-frequency components of reconstruction features into generation features, mitigating incompatibility while preserving the editing ability for major structures encoded in low-frequency information. Second, it introduces controlled noise into the replaced reconstruction features, expanding the generative space to enable diverse non-rigid edits beyond the original image’s constraints. Experiments on non-rigid edits, e.g. , addition, deletion, and pose manipulation, demonstrate that FSI-Edit outperforms existing baselines in target alignment, semantic fidelity and visual
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image EditingJiahui Sun, Weining Wang, Mingzhen Sun, Peiyao Wang 等ICLR 2026
- FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image EditingKaixiang Yang, Boyang Shen, Xin Li, Yuchen Dai 等AAAI 2026 · 被引用 3 次
- FeedEdit: Text-Based Image Editing with Dynamic Feedback RegulationFengyi Fu, Lei Zhang, Mengqi Huang, Zhendong MaoCVPR 2025
- DirectEdit: Step-Level Accurate Inversion for Flow-Based Image EditingDesong Yang, Mang YeICML 2026
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image EditingYucheng Liao, Jiajun Liang, Kaiqian Cui, Baoquan Zhao 等CVPR 2026 · 被引用 6 次
