: Sigmoid Modulation for Ultra High Resolution Diffusion
Bingxuan Zhao, Qing Zhou, Yu Wang, Chuang Yang, Qi Wang
摘要
Diffusion Transformers (DiTs) can synthesize high-fidelity images, but training at ultra-high resolutions is expensive, making inference-time extrapolation essential. Existing methods are typically scale-agnostic, applying the same positional-modulation schedule regardless of target resolution. We show that this misses a scale-sensitive property of denoising: resizing shifts a fixed semantic pattern toward lower normalized spatial frequencies while enlarging the spatial support over which global structure must be coordinated. The latter delays structural lock-in at high resolutions, so schedules may either relax guidance before global layout has stabilized, causing structural collapse, or retain excessive intervention into late denoising, causing textural degradation. We introduce SigMa (), a training-free framework that uses Sigmoid Modulation for scale-adaptive extrapolation through two scaling laws: Decoupled Geometric Center Alignment and Iso-Variance Rate Adaptation. Experiments show that SigMa reduces this mismatch, enabling training-free extrapolation up to 16 megapixels and achieving the best or competitive performance among the tested training-free extrapolation baselines. Code is available at github.com/bxuanz/SigMa.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution DiffusionNoam Issachar, Guy Yariv, Sagie Benaim, Yossi Adi 等ICML 2026 · 被引用 14 次
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion TransformersYiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao 等CVPR 2026 · 被引用 1 次
- FiT: Flexible Vision Transformer for Diffusion ModelZeyu Lu, Zidong Wang, Di Huang, Chengyue Wu 等ICML 2024 · 被引用 83 次
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic DiffusionSungho Koh, SeungJu Cha, Hyunwoo Oh, Kwanyoung Lee 等NeurIPS 2025 · 被引用 5 次
- LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion TransformerSong Fei, Tian Ye, Lujia Wang, Lei ZhuICLR 2026 · 被引用 9 次
