Transformers Can Learn Temporal Difference Methods for In-Context Reinforcement Learning
Jiuqi Wang, Ethan Blaser, Hadi Daneshmand, Shangtong Zhang
Abstract
Traditionally, reinforcement learning (RL) agents learn to solve new tasks by updating their neural network parameters through interactions with the task environment. However, recent works demonstrate that some RL agents, after certain pretraining procedures, can learn to solve unseen new tasks without parameter updates, a phenomenon known as in-context reinforcement learning (ICRL). The empirical success of ICRL is widely attributed to the hypothesis that the forward pass of the pretrained agent neural network implements an RL algorithm. In this paper, we support this hypothesis by showing, both empirically and theoretically, that when a transformer is trained for policy evaluation tasks, it can discover and learn to implement temporal difference learning in its forward pass.
- Equal contribution. The order is determined by tossing a fair coin. † Work performed while affiliated with MIT LIDS/Boston University.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb325531-daca-425e-9cc1-57f6bc530c04Cited by top-tier papers9
- Reward Is Enough: LLMs Are In-Context Reinforcement LearnersKefan Song, Amir Moeini, Peng Wang, Lei Gong et al.ICLR 2026 · 42 citations
- Towards Provable Emergence of In-Context Reinforcement LearningJiuqi Wang, Rohan Chandra, Shangtong ZhangNeurIPS 2025 · 5 citations
- Safe In-Context Reinforcement LearningAmir Moeini, Minjae Kwon, Alper Bozkurt, Yuichi Motai et al.ICML 2026 · 4 citations
- Linear Transformers Implicitly Discover Unified Numerical AlgorithmsPatrick Lutz, Aditya Gangrade, Hadi Daneshmand, Venkatesh SaligramaNeurIPS 2025 · 3 citations
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better InferenceYue Yu, Qiwei Di, Quanquan Gu, Dongruo ZhouICML 2026 · 3 citations
Builds on37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
Related papers
- Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised PretrainingLicong Lin, Yu Bai, Song MeiICLR 2024 · 74 citations
- In-context Exploration-Exploitation for Reinforcement LearningZhenwen Dai, Federico Tomasi, Sina GhiassianICLR 2024 · 14 citations
- Vintix II: Decision Pre-Trained Transformer is a Scalable In-Context Reinforcement LearnerAndrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Artyom Grishin et al.ICLR 2026
- On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse RecoveryRenpu Liu, Ruida Zhou, Cong Shen, Jing YangICLR 2025
- Distilling Reinforcement Learning Algorithms for In-Context Model-Based PlanningJaehyeon Son, Soochan Lee, Gunhee KimICLR 2025
