Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
Xinchen Yan, Chen Liang, Lijun Yu, Adams Wei Yu, Yifeng Lu, Quoc Le
Abstract
This paper investigates the scaling properties of autoregressive next-pixel prediction, a simple, end-to-end yet under-explored framework for unified vision models. Starting with images at resolutions of 32 32, we train a family of Transformers using IsoFLOP profiles across compute budgets up to 7e19 FLOPs and evaluate three distinct target metrics: next-pixel prediction objective, ImageNet classification accuracy, and generation-based completion where top half of the image serves as a spatial prompt. First, optimal scaling strategy is critically task-dependent. At a fixed resolution of 32 32 alone, the optimal scaling properties for image classification and image generation diverge, where generation optimal setup requires the data size grow three to five times faster than for the classification optimal setup. Second, as image resolution increases, the optimal scaling strategy indicates that the model size must grow much faster than data size. Surprisingly, by projecting our findings, we discover that the primary bottleneck is compute rather than the amount of training data. As compute continues to grow four to five times annually, we forecast the feasibility of pixel-by-pixel modeling of images within the next five years.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f52d7efa-1a35-4e6a-852d-06f4ff3b7e98Cited by top-tier papers1
Ask how each one uses itBuilds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Next Patch Prediction for AutoRegressive Visual GenerationYatian Pang, Peng Jin, Shuo Yang, Bin Zhu et al.AAAI 2026 · 24 citations
- FlowAR: Scale-wise Autoregressive Image Generation Meets Flow MatchingSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen et al.ICML 2025
- MoSs: Mixture of Scales for Efficient High-Resolution Autoregressive Image GenerationYaoxiu Lian, Hao Liang, Zhihong Gou, Yijia Zhang et al.AAAI 2026
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive TransformersBoxun Xu, Yu Wang, Zihu Wang, Peng LiAAAI 2026 · 2 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
