FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded Sampling
Zhengqiang Zhang, Ruihuang Li, Lei Zhang
Abstract
While image generation with diffusion models has achieved a great success, generating images of higher resolution than the training size remains a challenging task due to the high computational cost. Current methods typically perform the entire sampling process at full resolution and process all frequency components simultaneously, contradicting with the inherent coarse-to-fine nature of latent diffusion models and wasting computations on processing premature high-frequency details at early diffusion stages. To address this issue, we introduce an efficient Frequency-aware Cascaded Sampling framework, FreCaS in short, for higherresolution image generation. FreCaS decomposes the sampling process into cascaded stages with gradually increased resolutions, progressively expanding frequency bands and refining the corresponding details. We propose an innovative frequency-aware classifier-free guidance (FA-CFG) strategy to assign different guidance strengths for different frequency components, directing the diffusion model to add new details in the expanded frequency domain of each stage. Additionally, we fuse the cross-attention maps of previous and current stages to avoid synthesizing unfaithful layouts. Experiments demonstrate that FreCaS significantly outperforms state-of-the-art methods in image quality and generation speed. In particular, FreCaS is about 2.86× and 6.07× faster than ScaleCrafter and DemoFusion in generating a 2048×2048 image using a pre-trained SDXL model and achieves an FID b improvement of 11.6 and 3.7, respectively. FreCaS can be easily extended to more complex models such as SD3. The source code of FreCaS can be found at https://github.com/xtudbxk/FreCaS . INTRODUCATION In recent years, diffusion models, such as Imagen (Saharia et al., 2022) , SDXL (Podell et al., 2023) , PixelArt-α (Chen et al., 2023) and SD3 Esser et al. (2024), have achieved a remarkable success in generating high-quality natural images. However, these models face challenges in generating very high resolution images due to the increased complexity in high-dimensional space. Though efficient diffusion models, including ADM (Dhariwal & Nichol, 2021 ), CascadedDM (Ho et al., 2022) and LDM (Rombach et al., 2022) , have been developed, the computational burden of training diffusion models from scratch for high-resolution image generation remains substantial. As a result, popular diffusion models, such as SDXL (Podell et al., 2023) and SD3 (Esser et al., 2024), primarily focus on generating 1024 × 1024 resolution images. It is thus increasingly attractive to explore trainingfree strategies for generating images at higher resolutions, such as 2048 × 2048 and 4096 × 4096, using pre-trained diffusion models. MultiDiffusion (Bar-Tal et al., 2023) is among the first works to synthesize higher-resolution images using pre-trained diffusion models. However, it suffers from issues such as object duplication, which largely reduces the image quality. To address these issues, Jin et al. (2024) proposed to manually adjust the scale of entropy in the attention operations. He et al. (2023) and Huang et al. (2024)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cc4c267-c45c-4e0b-9ffd-b71e78bfaddaCited by top-tier papers7
- UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality DatasetChen Zhao, En Ci, Yunzhe Xu, Tiehan Fan et al.NeurIPS 2025 · 24 citations
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and GenerationZhengqiang Zhang, Rongyuan Wu, Lingchen Sun, Lei ZhangNeurIPS 2025 · 8 citations
- SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D EditingRuihuang Li, Liyi Chen, Zhengqiang Zhang, Varun Jampani et al.AAAI 2025 · 4 citations
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion TransformersYiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao et al.CVPR 2026 · 1 citation
- Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image GenerationNadav Z. Cohen, Ofir Abramovich, Ariel ShamirSIGGRAPH 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion ModelsYingqing He, Shaoshu Yang, Haoxin Chen, Xiaodong Cun et al.ICLR 2024 · 125 citations
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic DiffusionSungho Koh, SeungJu Cha, Hyunwoo Oh, Kwanyoung Lee et al.NeurIPS 2025 · 5 citations
- FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale FusionHaonan Qiu, Shiwei Zhang, Yujie Wei, Ruihang Chu et al.ICCV 2025 · 5 citations
- InstantAS: Minimum Coverage Sampling for Arbitrary-Size Image GenerationChangshuo Wang, Mingzhe Yu, Lei Wu, Lei Meng et al.ACM MM 2024
- simple diffusion: End-to-end diffusion for high resolution imagesEmiel Hoogeboom, Jonathan Heek, Tim SalimansICML 2023 · 403 citations
