MoSs: Mixture of Scales for Efficient High-Resolution Autoregressive Image Generation
Yaoxiu Lian, Hao Liang, Zhihong Gou, Yijia Zhang, Jiaming Xu, Guohao Dai, Ningyi Xu
Abstract
Since next-scale prediction was introduced as a new paradigm for autoregressive image generation, it has attracted extensive research interest. By progressively increasing resolution in a draft-to-refinement process, next-scale prediction demonstrates great potential in both generation quality and efficiency. However, at high resolutions, this paradigm faces a fundamental challenge: token sequences grow quadratically and accumulate across multiple scales, resulting in a critical performance bottleneck. Our systematic study reveals two key observations: (1) most image regions stabilize during early drafting stages, rendering subsequent full-scale refinement token-inefficient; and (2) different scales inherently present efficiency-fidelity trade-offs, suggesting that adaptive token dispatch across scales can concentrate computational resources where they yield the greatest quality gains. Motivated by these insights, we propose a training-free Mixture of Scales (MoSs) method for efficient high-resolution autoregressive image generation. MoSs breaks the strict causal dependency across scales in the final refinement steps by parallelizing scales of different resolutions, with each scale responsible for a subset of spatial regions. A lightweight frequency-based token dispatcher analyzes the drafted image and assigns regions to the appropriate scale. The outputs are then composited over the draft to produce the final high-resolution image. The scale-mixture method achieves substantial efficiency improvements with minimal impact on generation quality across various models. For instance, our implementation achieves 2.05-4.96× speedup on transformer backbone, up to 85.62% KV cache reduction, while incurring only 0.1-2.4% loss on GenEval quality metrics, as demonstrated on the state-of-the-art Infinity model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bbaae59-e8ea-45d7-befa-3caf46b9461dBuilds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- Next-Scale Autoregressive Models for Text-to-Motion GenerationZhiwei Zheng, Shibo Jin, Lingjie Liu, Mingmin ZhaoCVPR 2026 · 6 citations
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image SynthesisZhuokun Chen, Jugang Fan, Zhuowei Yu, Bohan Zhuang et al.ICCV 2025 · 3 citations
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive TransformersBoxun Xu, Yu Wang, Zihu Wang, Peng LiAAAI 2026 · 2 citations
- LazyVAR: Accelerating Visual Autoregressive Models via Scale-wise Token Pruning and Parallel Group DecodingRongge Mao, Chengqi Dong, S Kevin ZhouCVPR 2026
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache CompressionKunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng HwangNeurIPS 2025 · 23 citations
