Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
Qi Zhang, Tianqi Du, Haotian Huang, Yifei Wang, Yisen Wang
Abstract
In recent years, the rise of generative selfsupervised learning (SSL) paradigms has exhibited impressive performance across visual, language, and multi-modal domains. While the varied designs of generative SSL objectives lead to distinct properties in downstream tasks, a theoretical understanding of these differences remains largely unexplored. In this paper, we establish the first theoretical comparisons between two leading generative SSL paradigms: autoregressive SSL and masked SSL. Through establishing theoretical frameworks, we elucidate the strengths and limitations of autoregressive and masked SSL within the primary evaluation tasks of classification and content generation. Our findings demonstrate that in classification tasks, the flexibility of targeted tokens in masked SSL fosters more inter-sample connections compared to the fixed position of target tokens in autoregressive SSL, which yields superior clustering performance. In content generation tasks, the misalignment between the flexible lengths of test samples and the fixed length of unmasked texts in masked SSL (vs. flexible lengths of conditional texts in autoregressive SSL) hinders its generation performance. To leverage each other's strengths and mitigate weaknesses, we propose diversity-enhanced autoregressive and variable-length masked objectives, which substantially improve the classification performance of autoregressive SSL and the generation performance of masked SSL. Code is available at https://github.com/ PKU-ML/LookAheadLookAround .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 410b60a1-047d-4daa-8b7b-db7360ae7584Cited by top-tier papers2
- Long-Short Alignment for Effective Long-Context Modeling in LLMsTianqi Du, Haotian Huang, Yifei Wang, Yisen WangICML 2025
- Elucidating the design space of language models for image generationXuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu et al.ICML 2025
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
Related papers
- Understand Before You Generate: Self-Guided Training for Autoregressive Image GenerationXiaoyu Yue, Zidong Wang, Yuqing Wang, Wenlong Zhang et al.NeurIPS 2025 · 9 citations
- Reverse Engineering Self-Supervised LearningIdo Ben-Shaul, Ravid Shwartz-Ziv, Tomer Galanti, Shai Dekel et al.NeurIPS 2023 · 55 citations
- Can Generative Models Improve Self-Supervised Representation Learning?Sana Ayromlou, Vahid Reza Khazaie, Fereshteh Forghani, Arash AfkanpourAAAI 2025 · 5 citations
- Rethinking Graph Masked Autoencoders through Alignment and UniformityLiang Wang, Xiang Tao, Qiang Liu, Shu Wu et al.AAAI 2024 · 40 citations
- Cluster-Aware Contrastive Multi-View Clustering Based on Masked ViewsPenglei Wang, Ziming Quan, Danyang Wu, Jin XuACM MM 2025
