Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
Xiang Gao, Zhengbo Xu, Junhan Zhao, Jiaying Liu
Abstract
Recently, large-scale text-to-image (T2I) diffusion models have emerged as a powerful tool for image-to-image translation ( I2I ), allowing open-domain image translation via user-provided text prompts. This paper proposes frequencycontrolled diffusion model (FCDiffusion), an end-to-end diffusion-based framework that contributes a novel solution to text-guided I2I from a frequency-domain perspective. At the heart of our framework is a feature-space frequencydomain filtering module based on Discrete Cosine Transform, which filters the latent features of the source image in the DCT domain, yielding filtered image features bearing different DCT spectral bands as different control signals to the pre-trained Latent Diffusion Model. We reveal that control signals of different DCT spectral bands bridge the source image and the T2I generated image in different correlations (e.g., style, structure, layout, contour, etc.), and thus enable versatile I2I applications emphasizing different I2I correlations, including style-guided content creation, image semantic manipulation, image scene translation, and image style translation. Different from related approaches, FCDiffusion establishes a unified text-guided I2I framework suitable for diverse image translation tasks simply by switching among different frequency control branches at inference time. The effectiveness and superiority of our method for text-guided I2I are demonstrated with extensive experiments both qualitatively and quantitatively. Our project is publicly available at: https://xianggao1102.github.io/FCDiffusion/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 607c816c-a514-4cce-81b8-21743bb46070Cited by top-tier papers15
- One Stone with Two Birds: A Null-Text-Null Frequency-Aware Diffusion Models for Text-Guided Image InpaintingHaipeng Liu, Yang Wang, Meng WangNeurIPS 2025 · 8 citations
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image EditingYucheng Liao, Jiajun Liang, Kaiqian Cui, Baoquan Zhao et al.CVPR 2026 · 6 citations
- FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image TranslationXiang Gao, Jiaying LiuACM MM 2024 · 5 citations
- FRBNet: Revisiting Low-Light Vision through Frequency-Domain Radial Basis NetworkFangtong Sun, Congyu Li, Ke Yang, Yuchen Pan et al.NeurIPS 2025 · 4 citations
- Exploring the Design Space of Diffusion Bridge ModelsShaorong Zhang, Yuanbin Cheng, Greg Ver SteegNeurIPS 2025 · 3 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Frequency-Guided Diffusion for Training-Free Text-Driven Image TranslationZheng Gao, Jifei Song, Zhensong Zhang, Jiankang Deng et al.ICCV 2025 · 1 citation
- ControlStyle: Text-Driven Stylized Image Generation Using Diffusion PriorsJingwen Chen, Yingwei Pan, Ting Yao, Tao MeiACM MM 2023 · 45 citations
- Frequency Domain Image Translation: More Photo-realistic, Better Identity-preservingMu Cai, Hong Zhang, Huijuan Huang, Qichuan Geng et al.ICCV 2021 · 118 citations
- FSI-Edit: Frequency and Stochasticity Injection for Flexible Diffusion-Based Image EditingKaixiang Yang, Xin Li, Yuxi Li, Qiang Li et al.NeurIPS 2025 · 2 citations
- Stylekeeper: Prevent Content Leakage using Negative Visual Query GuidanceJaeseok Jeong, Junho Kim, Gayoung Lee, Yunjey Choi et al.ICCV 2025
