FAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusion
Haosen Yang, Adrian Bulat, Isma Hadji, Hai X. Pham, Xiatian Zhu, Georgios Tzimiropoulos, Brais Martínez
Abstract
Diffusion models are proficient at generating high-quality images. They are however effective only when operating at the resolution used during training. Inference at a scaled resolution leads to repetitive patterns and structural distortions. Retraining at higher resolutions quickly becomes prohibitive. Thus, methods enabling pre-existing diffusion models to operate at flexible test-time resolutions are highly desirable. Previous works suffer from frequent artifacts and often introduce large latency overheads. We propose two simple modules that combine to solve these issues. We introduce a Frequency Modulation (FM) module that leverages the Fourier domain to improve the global structure consistency, and an Attention Modulation (AM) module which improves the consistency of local texture patterns, a problem largely ignored in prior works. Our method, coined FAM diffusion, can seamlessly integrate into any latent diffusion model and requires no additional training. Extensive qualitative results highlight the effectiveness of our method in addressing structural and local artifacts, while quantitative results show state-of-the-art performance. Also, our method avoids redundant inference tricks for improved consistency such as patch-based or progressive generation, leading to negligible latency overheads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution DiffusionNoam Issachar, Guy Yariv, Sagie Benaim, Yossi Adi et al.ICML 2026 · 14 citations
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion TransformersYiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao et al.CVPR 2026 · 1 citation
- LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance FlowYuan Zhou, Yan Zhang, Jianlong Chang, Xin Gu et al.AAAI 2026
- Exploring Position Encoding Mechanism in Diffusion U-Net for Training-free High-resolution Image GenerationFeng Zhou, Pu Cao, Yiyang Ma, Lu Yang et al.AAAI 2026
- FreeAdapt: Unleashing Diffusion Priors for Ultra-High-Definition Image RestorationXiaoan Liu, Xinyi Liu, Yongjun Zhang, Yi Wan et al.ICLR 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic DiffusionSungho Koh, SeungJu Cha, Hyunwoo Oh, Kwanyoung Lee et al.NeurIPS 2025 · 5 citations
- AID: Attention Interpolation of Text-to-Image DiffusionQiyuan He, Jinghao Wang, Ziwei Liu, Angela YaoNeurIPS 2024 · 30 citations
- GenesisTex2: Stable, Consistent and High-Quality Text-to-Texture GenerationJiawei Lu, Yingpeng Zhang, Zengjun Zhao, He Wang et al.AAAI 2025 · 10 citations
- Frequency-Aware Flow Matching for High-Quality Image GenerationSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen et al.CVPR 2026 · 6 citations
- FreeControl: Efficient, Training-Free Structural Control via One-Step Attention ExtractionJiang Lin, Xinyu Chen, Song Wu, Zhiqiu Zhang et al.NeurIPS 2025 · 3 citations
