Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
Xiang Gao, Zhengbo Xu, Junhan Zhao, Jiaying Liu
摘要
Recently, large-scale text-to-image (T2I) diffusion models have emerged as a powerful tool for image-to-image translation ( I2I ), allowing open-domain image translation via user-provided text prompts. This paper proposes frequencycontrolled diffusion model (FCDiffusion), an end-to-end diffusion-based framework that contributes a novel solution to text-guided I2I from a frequency-domain perspective. At the heart of our framework is a feature-space frequencydomain filtering module based on Discrete Cosine Transform, which filters the latent features of the source image in the DCT domain, yielding filtered image features bearing different DCT spectral bands as different control signals to the pre-trained Latent Diffusion Model. We reveal that control signals of different DCT spectral bands bridge the source image and the T2I generated image in different correlations (e.g., style, structure, layout, contour, etc.), and thus enable versatile I2I applications emphasizing different I2I correlations, including style-guided content creation, image semantic manipulation, image scene translation, and image style translation. Different from related approaches, FCDiffusion establishes a unified text-guided I2I framework suitable for diverse image translation tasks simply by switching among different frequency control branches at inference time. The effectiveness and superiority of our method for text-guided I2I are demonstrated with extensive experiments both qualitatively and quantitatively. Our project is publicly available at: https://xianggao1102.github.io/FCDiffusion/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- One Stone with Two Birds: A Null-Text-Null Frequency-Aware Diffusion Models for Text-Guided Image InpaintingHaipeng Liu, Yang Wang, Meng WangNeurIPS 2025 · 被引用 8 次
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image EditingYucheng Liao, Jiajun Liang, Kaiqian Cui, Baoquan Zhao 等CVPR 2026 · 被引用 6 次
- FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image TranslationXiang Gao, Jiaying LiuACM MM 2024 · 被引用 5 次
- FRBNet: Revisiting Low-Light Vision through Frequency-Domain Radial Basis NetworkFangtong Sun, Congyu Li, Ke Yang, Yuchen Pan 等NeurIPS 2025 · 被引用 4 次
- Exploring the Design Space of Diffusion Bridge ModelsShaorong Zhang, Yuanbin Cheng, Greg Ver SteegNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Frequency-Guided Diffusion for Training-Free Text-Driven Image TranslationZheng Gao, Jifei Song, Zhensong Zhang, Jiankang Deng 等ICCV 2025 · 被引用 1 次
- ControlStyle: Text-Driven Stylized Image Generation Using Diffusion PriorsJingwen Chen, Yingwei Pan, Ting Yao, Tao MeiACM MM 2023 · 被引用 45 次
- Frequency Domain Image Translation: More Photo-realistic, Better Identity-preservingMu Cai, Hong Zhang, Huijuan Huang, Qichuan Geng 等ICCV 2021 · 被引用 118 次
- FSI-Edit: Frequency and Stochasticity Injection for Flexible Diffusion-Based Image EditingKaixiang Yang, Xin Li, Yuxi Li, Qiang Li 等NeurIPS 2025 · 被引用 2 次
- Stylekeeper: Prevent Content Leakage using Negative Visual Query GuidanceJaeseok Jeong, Junho Kim, Gayoung Lee, Yunjey Choi 等ICCV 2025
