MoSs: Mixture of Scales for Efficient High-Resolution Autoregressive Image Generation
Yaoxiu Lian, Hao Liang, Zhihong Gou, Yijia Zhang, Jiaming Xu, Guohao Dai, Ningyi Xu
摘要
Since next-scale prediction was introduced as a new paradigm for autoregressive image generation, it has attracted extensive research interest. By progressively increasing resolution in a draft-to-refinement process, next-scale prediction demonstrates great potential in both generation quality and efficiency. However, at high resolutions, this paradigm faces a fundamental challenge: token sequences grow quadratically and accumulate across multiple scales, resulting in a critical performance bottleneck. Our systematic study reveals two key observations: (1) most image regions stabilize during early drafting stages, rendering subsequent full-scale refinement token-inefficient; and (2) different scales inherently present efficiency-fidelity trade-offs, suggesting that adaptive token dispatch across scales can concentrate computational resources where they yield the greatest quality gains. Motivated by these insights, we propose a training-free Mixture of Scales (MoSs) method for efficient high-resolution autoregressive image generation. MoSs breaks the strict causal dependency across scales in the final refinement steps by parallelizing scales of different resolutions, with each scale responsible for a subset of spatial regions. A lightweight frequency-based token dispatcher analyzes the drafted image and assigns regions to the appropriate scale. The outputs are then composited over the draft to produce the final high-resolution image. The scale-mixture method achieves substantial efficiency improvements with minimal impact on generation quality across various models. For instance, our implementation achieves 2.05-4.96× speedup on transformer backbone, up to 85.62% KV cache reduction, while incurring only 0.1-2.4% loss on GenEval quality metrics, as demonstrated on the state-of-the-art Infinity model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
相关 Paper
- Next-Scale Autoregressive Models for Text-to-Motion GenerationZhiwei Zheng, Shibo Jin, Lingjie Liu, Mingmin ZhaoCVPR 2026 · 被引用 6 次
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image SynthesisZhuokun Chen, Jugang Fan, Zhuowei Yu, Bohan Zhuang 等ICCV 2025 · 被引用 3 次
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive TransformersBoxun Xu, Yu Wang, Zihu Wang, Peng LiAAAI 2026 · 被引用 2 次
- LazyVAR: Accelerating Visual Autoregressive Models via Scale-wise Token Pruning and Parallel Group DecodingRongge Mao, Chengqi Dong, S Kevin ZhouCVPR 2026
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache CompressionKunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng HwangNeurIPS 2025 · 被引用 23 次
