: Sigmoid Modulation for Ultra High Resolution Diffusion
Bingxuan Zhao, Qing Zhou, Yu Wang, Chuang Yang, Qi Wang
Abstract
Diffusion Transformers (DiTs) can synthesize high-fidelity images, but training at ultra-high resolutions is expensive, making inference-time extrapolation essential. Existing methods are typically scale-agnostic, applying the same positional-modulation schedule regardless of target resolution. We show that this misses a scale-sensitive property of denoising: resizing shifts a fixed semantic pattern toward lower normalized spatial frequencies while enlarging the spatial support over which global structure must be coordinated. The latter delays structural lock-in at high resolutions, so schedules may either relax guidance before global layout has stabilized, causing structural collapse, or retain excessive intervention into late denoising, causing textural degradation. We introduce SigMa (), a training-free framework that uses Sigmoid Modulation for scale-adaptive extrapolation through two scaling laws: Decoupled Geometric Center Alignment and Iso-Variance Rate Adaptation. Experiments show that SigMa reduces this mismatch, enabling training-free extrapolation up to 16 megapixels and achieving the best or competitive performance among the tested training-free extrapolation baselines. Code is available at github.com/bxuanz/SigMa.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af3145c7-04a5-49c9-995c-c66bc2d1d932Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution DiffusionNoam Issachar, Guy Yariv, Sagie Benaim, Yossi Adi et al.ICML 2026 · 14 citations
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion TransformersYiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao et al.CVPR 2026 · 1 citation
- FiT: Flexible Vision Transformer for Diffusion ModelZeyu Lu, Zidong Wang, Di Huang, Chengyue Wu et al.ICML 2024 · 83 citations
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic DiffusionSungho Koh, SeungJu Cha, Hyunwoo Oh, Kwanyoung Lee et al.NeurIPS 2025 · 5 citations
- LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion TransformerSong Fei, Tian Ye, Lujia Wang, Lei ZhuICLR 2026 · 9 citations
