Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
Mathieu Blondel, Michael Sander, Germain Vivier-Ardisson, Tianlin Liu, Vincent Roulet
摘要
Autoregressive models (ARMs) currently constitute the dominant paradigm for large language models (LLMs). Energy-based models (EBMs) represent another class of models, which have historically been less prevalent in LLM development, yet naturally characterize the optimal policy in post-training alignment. In this paper, we present a unified view of these two model classes. Taking the chain rule of probability as a starting point, we establish an explicit bijection between ARMs and EBMs in function space, which we show to correspond to a special case of the soft Bellman equation in maximum entropy reinforcement learning. Building upon this bijection, we derive the equivalence between supervised learning of ARMs and EBMs. Furthermore, we analyze the distillation of EBMs into ARMs by providing theoretical error bounds. Our results provide insights into the ability of ARMs to plan ahead, despite being based on the next-token prediction paradigm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup 等NeurIPS 2021 · 被引用 565 次
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
- Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence ConstraintsChaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu 等ICLR 2024 · 被引用 173 次
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 被引用 163 次
相关 Paper
- Large Language Models Do Multi-Label Classification DifferentlyMarcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason 等EMNLP 2025
- Semantic-aware Wasserstein Policy Regularization for Large Language Model AlignmentByeonghu Na, Hyungho Na, Yeongmin Kim, Suhyeon Jo 等ICLR 2026 · 被引用 2 次
- Emo: Earth Mover Distance Optimization for Auto-Regressive Language ModelingSiyu Ren, Zhiyong Wu, Kenny Q. ZhuICLR 2024 · 被引用 9 次
- Enhancing Numerical Prediction of MLLMS With Soft LabelingPei Wang, Zhaowei Cai, Hao Yang, Davide Modolo 等ICCV 2025 · 被引用 1 次
- AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationSongming Zhang, Xue Zhang, Tong Zhang, Bojie Hu 等ACL 2025
