Frequency-Guided Diffusion for Training-Free Text-Driven Image Translation
Zheng Gao, Jifei Song, Zhensong Zhang, Jiankang Deng, Ioannis Patras
Abstract
Current training-free text-driven image translation primarily uses diffusion features (convolution and attention) of pre-trained model as guidance to preserve the style/structure of source image in translated image. However, the coarse guidance at feature level struggles with style (e.g., visual patterns) and structure (e.g., edges) alignment with the source. Based on the observation that the low-/high-frequency components retain style/structure information of image, in this work, we propose training-free Frequency-Guided Diffusion (FGD), which tailors low-/high-frequency guidance for style-and structure-guided translation, respectively. For low-frequency guidance (style-guided), we substitute the low-frequency components of diffusion latents from sampling process with those from inversion of source and normalize the obtained latent with composited spectrum to enforce color alignment. For highfrequency guidance (structure-guided), we propose highfrequency alignment and high-frequency injection that compensate each other. High-frequency alignment preserves edges and contour by adjusting the predicted noise with guidance function that aligns high-frequency image regions between sampling and source image. High-frequency injection facilitates layout preservation by injecting highfrequency components of diffusion convolution features (from inversion) to sampling process. Qualitative and quantitative results verify the superiority of our method on styleand structure-guided translation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 974421a0-18e7-4a56-bc7d-2263611d9cc0Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- StyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style TransferRuojun Xu, Weijie Xi, Xiaodi Wang, Yongbo Mao et al.CVPR 2025
- Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image TranslationXiang Gao, Zhengbo Xu, Junhan Zhao, Jiaying LiuAAAI 2024 · 23 citations
- Frequency Domain Image Translation: More Photo-realistic, Better Identity-preservingMu Cai, Hong Zhang, Huijuan Huang, Qichuan Geng et al.ICCV 2021 · 118 citations
- Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image GenerationNadav Z. Cohen, Ofir Abramovich, Ariel ShamirSIGGRAPH 2026
- W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image EditingJiahui Sun, Weining Wang, Mingzhen Sun, Peiyao Wang et al.ICLR 2026
