Denoising Autoregressive Representation Learning
Yazhe Li, Jörg Bornschein, Ting Chen
Abstract
In this paper, we explore a new generative approach for learning visual representations. Our method, DARL, employs a decoder-only Transformer to predict image patches autoregressively. We find that training with Mean Squared Error (MSE) alone leads to strong representations. To enhance the image generation ability, we replace the MSE loss with the diffusion objective by using a denoising patch decoder. We show that the learned representation can be improved by using tailored noise schedules and longer training in larger models. Notably, the optimal schedule differs significantly from the typical ones used in standard image diffusion models. Overall, despite its simple architecture, DARL delivers performance remarkably close to state-of-the-art masked prediction models under the fine-tuning protocol. This marks an important step towards a unified model capable of both visual perception and generation, effectively combining the strengths of autoregressive and denoising diffusion models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Latent Denoising Makes Good TokenizersJiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian et al.ICLR 2026 · 17 citations
- TabNAT: A Continuous-Discrete Joint Generative Framework for Tabular DataHengrui Zhang, Liancheng Fang, Qitian Wu, Philip S. YuICML 2025
- TimeDART: A Diffusion Autoregressive Transformer for Self-Supervised Time Series RepresentationDaoyu Wang, Mingyue Cheng, Zhiding Liu, Qi LiuICML 2025
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
Related papers
- Diffusion Models as Masked AutoencodersChen Wei, Karttikeya Mangalam, Po-Yao Huang, Yanghao Li et al.ICCV 2023 · 82 citations
- Denoising Autoregressive Transformers for Scalable Text-to-Image GenerationJiatao Gu, Yuyang Wang, Yizhe Zhang, Qihang Zhang et al.ICLR 2025
- D-AR: Diffusion via Autoregressive ModelsZiteng Gao, Mike Zheng ShouICLR 2026 · 11 citations
- Denoising Token Prediction in Masked Autoregressive ModelsTing Yao, Yehao Li, Yingwei Pan, Zhaofan Qiu et al.ICCV 2025 · 2 citations
- Self-Supervised Flow Matching for Scalable Multi-Modal SynthesisHila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell et al.ICML 2026 · 13 citations
