Sample- and Parameter-Efficient Auto-Regressive Image Models
Elad Amrani, Leonid Karlinsky, Alex M. Bronstein
2025Year
1Top-tier citations
Abstract
https://github.com/elad-amrani/xtra Figure 1. Sample and Parameter Efficiency of XTRA. (Left) XTRA-H/14 (0.6B parameters) outperforms prior state-of-the-art auto-regressive image model (AIM-0.6B [26]) in top-1 average accuracy across 15 diverse image recognition benchmarks, despite being trained on 152× fewer samples. (Right) XTRA-B/16 (85M parameters) outperforms prior auto-regressive image models trained on ImageNet-1k in linear and attentive probing tasks, while using 7-16× fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- HART: Efficient Visual Generation with Hybrid Autoregressive TransformerHaotian Tang, Yecheng Wu, Shang Yang, Enze Xie et al.ICLR 2025
- Autoregressive Pretraining with Mamba in VisionSucheng Ren, Xianhang Li, Haoqin Tu, Feng Wang et al.ICLR 2025
- FlexVAR: Flexible Visual Autoregressive Modeling without Residual PredictionSiyu Jiao, Gengwei Zhang, Yinlong Qian, Jiancheng Huang et al.NeurIPS 2025 · 23 citations
- Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on ImagesRewon ChildICLR 2021 · 45 citations
- RepVGG: Making VGG-Style ConvNets Great AgainXiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han et al.CVPR 2021
