NRGPT: An Energy-based Alternative for GPT
Nima Dehmamy, Benjamin Hoover, Bishwajit Saha, Leo Kozachkov, Jean-Jacques Slotine, Dmitry Krotov
摘要
Generative Pre-trained Transformer (GPT) architectures are the most popular design for language modeling. Energy-based modeling is a different paradigm that views inference as a dynamical process operating on an energy landscape. We propose a minimal modification of the GPT setting to unify it with the EBM framework. The inference step of our model, which we call eNeRgy-GPT (NRGPT), is conceptualized as an exploration of the tokens on the energy landscape. We prove, and verify empirically, that under certain circumstances this exploration becomes gradient descent, although they don't necessarily lead to the best performing models. We demonstrate that our model performs well for simple language (Shakespeare dataset), algebraic ListOPS tasks, and richer settings such as OpenWebText language modeling. We also observe that our models may be more resistant to overfitting, doing so only during very long training. Transformers represent a dominant paradigm in autoregressive language modeling (Vaswani et al., 2017) . In a typical setting, a sequence of tokens describing a text is passed through several transformer layers and mapped onto a new sequence, which is a copy of the original one shifted by one token and appended by the token that follows the initial sequence. At training time, this network is trained through self-supervised training, and at inference time the network is used for next token prediction. This is the standard Generative Pre-trained Transformer (GPT) setting, which is the first step in Large Language Model (LLM) design (Radford et al., 2018) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
相关 Paper
- GraphGPT: Generative Pre-trained Graph Eulerian TransformerQifang Zhao, Weidong Ren, Tianyu Li, Hong Liu 等ICML 2025
- Energy-Based Diffusion Language Models for Text GenerationMinkai Xu, Tomas Geffner, Karsten Kreis, Weili Nie 等ICLR 2025
- Residual Energy-Based Models for Text GenerationYuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam 等ICLR 2020 · 被引用 147 次
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token PredictionMathieu Blondel, Michael Sander, Germain Vivier-Ardisson, Tianlin Liu 等ICML 2026 · 被引用 8 次
- Transformers from an Optimization PerspectiveYongyi Yang, Zengfeng Huang, David P. WipfNeurIPS 2022 · 被引用 44 次
