Scalable Pre-training of Large Autoregressive Image Models
Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Ángel Bautista, Vaishaal Shankar, Alexander T. Toshev, Joshua M. Susskind, Armand Joulin
Abstract
This paper introduces AIM, a collection of vision models pre-trained with an autoregressive objective. These models are inspired by their textual counterparts, i.e., Large Language Models (LLMs), and exhibit similar scaling properties. Specifically, we highlight two key findings: (1) the performance of the visual features scale with both the model capacity and the quantity of data, (2) the value of the objective function correlates with the performance of the model on downstream tasks. We illustrate the practical implication of these findings by pre-training a 7 billion parameter AIM on 2 billion images, that achieves 84.0% on ImageNet-1k with a frozen trunk. Interestingly, even at this scale, we observe no sign of saturation in performance, suggesting that AIM potentially represents a new frontier for training large-scale vision models. The pre-training of AIM is similar to the pre-training of LLMs, and does not require any image-specific strategy to stabilize the training at scale.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f3a0f38-e30e-4198-be8d-84229b1bd09cCited by top-tier papers53
- Perception Encoder: The best visual embeddings are not at the output of the networkDaniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho et al.NeurIPS 2025 · 359 citations
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2024 · 331 citations
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 208 citations
- Unveiling Encoder-Free Vision-Language ModelsHaiwen Diao, Yufeng Cui, Xiaotong Li, Yueze Wang et al.NeurIPS 2024 · 107 citations
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and ModalitiesRoman Bachmann, Oguzhan Fatih Kar, David Mizrahi, Ali Garjani et al.NeurIPS 2024 · 60 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- Multimodal Autoregressive Pre-training of Large Vision EncodersEnrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter et al.CVPR 2025
- Elucidating the design space of language models for image generationXuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu et al.ICML 2025
- Sequential Modeling Enables Scalable Learning for Large Vision ModelsYutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar et al.CVPR 2024
- Multi-modal Auto-regressive Modeling via Visual TokensTianshuo Peng, Zuchao Li, Lefei Zhang, Hai Zhao et al.ACM MM 2024 · 1 citation
- Autoregressive Pretraining with Mamba in VisionSucheng Ren, Xianhang Li, Haoqin Tu, Feng Wang et al.ICLR 2025
