FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image Generation
Bonan Li, Yinhan Hu, Songhua Liu, Zeyu Xiao, Xinchao Wang
Abstract
Layout-to-Image generation has significantly advanced content creation by enabling the rendering of visual text under predefined spatial layouts. Current approaches achieve training-free layout guidance by constructing attention-based energy functions to derive correction gradients. In this paper, we demonstrate that vanilla energy functions suffer from two limitations, resulting in imprecise layout control and visually unrealistic artifacts. First, the normalizing factor of the Boltzmann distribution defined by the energy functions is non-negligible when calculating correction gradients, yet current energy functions cannot compute this factor exactly. Furthermore, while attention varies over time during the denoising process, existing approaches employ a fixed formulation. To address these challenges, we introduce FreLay, a novel training-free approach equipped with a frequency-aware energy function. Our method first reformulates the energy function to handle the normalization factor, enabling accurate computation of correction gradients. Simultaneously, leveraging the prior knowledge that low-frequency information deteriorates slower during noise addition, we design a time-specific energy function for each timestep from a frequency-domain perspective. Experimental results demonstrate that FreLay consistently outperforms existing state-of-the-art training-free methods by a large margin both qualitatively and quantitatively across multiple datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Control and Realism: Best of Both Worlds in Layout-to-Image without TrainingBonan Li, Yinhan Hu, Songhua Liu, Xinchao WangICML 2025
- Frequency-Guided Diffusion for Training-Free Text-Driven Image TranslationZheng Gao, Jifei Song, Zhensong Zhang, Jiankang Deng et al.ICCV 2025 · 1 citation
- Mitigating Noise-Induced Layout Priors for Object Counting in Diffusion ModelsXiaoling Gu, Xuelong Li, Shengqi Wu, Yongkang Wong et al.ICML 2026
- LoCo: Training-Free Layout-to-Image Synthesis with Localized ConstraintsPeiang Zhao, Han Li, Ruiyang Jin, S. Kevin ZhouACM MM 2025 · 2 citations
- W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image EditingJiahui Sun, Weining Wang, Mingzhen Sun, Peiyao Wang et al.ICLR 2026
