Next-Token Prediction and Regret Minimization
Mehryar Mohri, Clayton Sanford, Jon Schneider, Kiran Vodrahalli, Yifan Wu
摘要
We consider the question of how to employ next-token prediction algorithms in adversarial online decision making environments. Specifically, if we train a next-token prediction model on a distribution D over sequences of opponent actions, when is it the case that the induced online decision making algorithm (by approximately best responding to the model's predictions) has low adversarial regret (i.e., when is D a low-regret distribution)? For unbounded context windows (where the prediction made by the model can depend on all the actions taken by the adversary thus far), we show that although not every distribution D is a low-regret distribution, every distribution D is exponentially close (in TV distance) to one lowregret distribution, and hence sublinear regret can always be achieved at negligible cost to the accuracy of the original next-token prediction model. In contrast to this, for bounded context windows (where the prediction made by the model can depend only on the past w actions taken by the adversary, as may be the case in modern transformer architectures), we show that there are some distributions D of opponent play that are Θ(1)-far from any low-regret distribution D ′ (even when w = Ω(T ) and such distributions exist). Finally, we complement these results by showing that the unbounded context robustification procedure can be implemented by layers of a standard transformer architecture, and provide empirical evidence that transformer models can be efficiently trained to represent these new low-regret distributions. * Work done when Yifan Wu was a student researcher with Google Research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Representational Strengths and Limitations of TransformersClayton Sanford, Daniel J. Hsu, Matus TelgarskyNeurIPS 2023 · 被引用 162 次
- Can large language models explore in-context?Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang 等NeurIPS 2024 · 被引用 95 次
- Transformers, parallel computation, and logarithmic depthClayton Sanford, Daniel Hsu, Matus TelgarskyICML 2024 · 被引用 64 次
- Online Learning with Bounded RecallJon Schneider, Kiran VodrahalliICML 2024 · 被引用 1 次
相关 Paper
- Adversarially Robust Decision TransformerXiaohang Tang, Afonso Marques, Parameswaran Kamalaruban, Ilija BogunovicNeurIPS 2024 · 被引用 5 次
- Adversarially Pretrained Transformers May Be Universally Robust In-Context LearnersSoichiro Kumano, Hiroshi Kera, Toshihiko YamasakiICLR 2026 · 被引用 2 次
- Robust In-Context Reinforcement Learning Under Reward Poisoning AttacksPaulius Sasnauskas, Yiğit Yalın, Goran RadanovicICML 2026
- Online and Distribution-Free Robustness: Regression and Contextual Bandits with Huber ContaminationSitan Chen, Frederic Koehler, Ankur Moitra, Morris YauFOCS 2021 · 被引用 14 次
- Near-Optimal Online Learning with Non-Stochastic and Unbounded Erroneous FeedbackDacheng Wen, Yupeng Li, Francis C. M. Lau, Tian Wang 等INFOCOM 2026
