Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
Mathieu Blondel, Michael Sander, Germain Vivier-Ardisson, Tianlin Liu, Vincent Roulet
Abstract
Autoregressive models (ARMs) currently constitute the dominant paradigm for large language models (LLMs). Energy-based models (EBMs) represent another class of models, which have historically been less prevalent in LLM development, yet naturally characterize the optimal policy in post-training alignment. In this paper, we present a unified view of these two model classes. Taking the chain rule of probability as a starting point, we establish an explicit bijection between ARMs and EBMs in function space, which we show to correspond to a special case of the soft Bellman equation in maximum entropy reinforcement learning. Building upon this bijection, we derive the equivalence between supervised learning of ARMs and EBMs. Furthermore, we analyze the distillation of EBMs into ARMs by providing theoretical error bounds. Our results provide insights into the ability of ARMs to plan ahead, despite being based on the next-token prediction paradigm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d0b922d-93ed-4f73-a08b-6e452262c8beCited by top-tier papers1
Ask how each one uses itBuilds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun et al.NeurIPS 2022 · 316 citations
- Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence ConstraintsChaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu et al.ICLR 2024 · 173 citations
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 163 citations
Related papers
- Large Language Models Do Multi-Label Classification DifferentlyMarcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason et al.EMNLP 2025
- Semantic-aware Wasserstein Policy Regularization for Large Language Model AlignmentByeonghu Na, Hyungho Na, Yeongmin Kim, Suhyeon Jo et al.ICLR 2026 · 2 citations
- Emo: Earth Mover Distance Optimization for Auto-Regressive Language ModelingSiyu Ren, Zhiyong Wu, Kenny Q. ZhuICLR 2024 · 9 citations
- Enhancing Numerical Prediction of MLLMS With Soft LabelingPei Wang, Zhaowei Cai, Hao Yang, Davide Modolo et al.ICCV 2025 · 1 citation
- AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationSongming Zhang, Xue Zhang, Tong Zhang, Bojie Hu et al.ACL 2025
