ICML2026
: Sigmoid Modulation for Ultra High Resolution Diffusion
Bingxuan Zhao, Qing Zhou, Yu Wang, Chuang Yang, Qi Wang
摘要
Diffusion Transformers (DiTs) can synthesize high-fidelity images, but training at ultra-high resolutions is expensive, making inference-time extrapolation essential. Existing methods are typically scale-agnostic, applying the same positional-modulation schedule regardless of target resolution. We show that this misses a scale-sensitive property of denoising: resizing shifts a fixed semantic pattern toward lower normalized spatial frequencies while enlarging the spatial support over which global structure must be coordinated. The latter delays structural lock-in at high resolutions, so schedules may either relax guidance before global layout has stabilized, causing structural collapse, or retain excessive intervention into late denoising, causing textural degradation. We introduce SigMa (), a training-free framework that uses Sigmoid Modulation for scale-adaptive extrapolation through two scaling laws: Decoupled Geometric Center Alignment and Iso-Variance Rate Adaptation. Experiments show that SigMa reduces this mismatch, enabling training-free extrapolation up to 16 megapixels and achieving the best or competitive performance among the tested training-free extrapolation baselines. Code is available at github.com/bxuanz/SigMa.