Masked Generative Nested Transformers with Decode Time Scaling
Sahil Goyal, Debapriya Tula, Gagan Jain, Pradeep Shenoy, Prateek Jain, Sujoy Paul
Abstract
Recent advances in visual generation have made significant strides in producing content of exceptional quality. However, most methods suffer from a fundamental problem -a bottleneck of inference computational efficiency. Most of these algorithms involve multiple passes over a transformer model to generate tokens or denoise inputs. However, the model size is kept consistent for all iterations, making it computationally expensive. In this work, we aim to address this issue primarily through two key ideas -(a) not all parts of the generation process need equal compute, hence we design a decode time model scaling schedule to utilize compute effectively, and (b) we can cache and reuse some of the intermediate computation. Combining these two ideas leads to using smaller models to process more tokens while large models process fewer tokens. These different-sized models do not increase the parameter size, as they share parameters. We rigorously experiment with ImageNet256×256 , ImageNet128×128, UCF101, and Kinetics600 to showcase the efficacy of the proposed method for image/video generation and frame prediction. Our experiments show that with almost 3× less compute than baseline, our model obtains competitive performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on58
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- DDiT: Dynamic Patch Scheduling for Efficient Diffusion TransformersDahye Kim, Deepti Ghadiyaram, Raghudeep GaddeCVPR 2026 · 3 citations
- Learnings from Scaling Visual Tokenizers for Reconstruction and GenerationPhilippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar et al.ICML 2025
- ScalingCache: Extreme Acceleration of DiTs through Difference Scaling and Dynamic Interval CachingLihui Gu, Jingbin He, Lianghao Su, Kang He et al.ICLR 2026
- Towards Precise Scaling Laws for Video Diffusion TransformersYuanyang Yin, Yaqi Zhao, Mingwu Zheng, Ke Lin et al.CVPR 2025
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive GenerationZhekai Chen, Ruihang Chu, Yukang Chen, Shiwei Zhang et al.NeurIPS 2025 · 16 citations
