In-context Reinforcement Learning with Algorithm Distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Stenberg Hansen, Angelos Filos, Ethan Brooks, Maxime Gazeau, Himanshu Sahni
Abstract
We propose Algorithm Distillation (AD), a method for distilling reinforcement learning (RL) algorithms into neural networks by modeling their training histories with a causal sequence model. Algorithm Distillation treats learning to reinforcement learn as an across-episode sequential prediction problem. A dataset of learning histories is generated by a source RL algorithm, and then a causal transformer is trained by autoregressively predicting actions given their preceding learning histories as context. Unlike sequential policy prediction architectures that distill post-learning or expert sequences, AD is able to improve its policy entirely in-context without updating its network parameters. We demonstrate that AD can reinforcement learn in-context in a variety of environments with sparse rewards, combinatorial task structure, and pixel-based observations, and find that AD learns a more data-efficient RL algorithm than the one that generated the source data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b61505b-d17e-4666-9a85-427ea2e1b4e4Cited by top-tier papers50
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt et al.ICLR 2024 · 243 citations
- The Learnability of In-Context LearningNoam Wies, Yoav Levine, Amnon ShashuaNeurIPS 2023 · 207 citations
- GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local RefinementsAlexander Havrilla, Sharath Chandra Raparthy, Christoforos Nalmpantis, Jane Dwivedi-Yu et al.ICML 2024 · 105 citations
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma et al.ICLR 2023 · 85 citations
- Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence ModelingSili Huang, Jifeng Hu, Zhejian Yang, Liwei Yang et al.NeurIPS 2024 · 65 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang et al.NeurIPS 2022 · 407 citations
Related papers
- Efficient Transformers in Reinforcement Learning using Actor-Learner DistillationEmilio Parisotto, Ruslan SalakhutdinovICLR 2021 · 51 citations
- Distilling Reinforcement Learning Algorithms for In-Context Model-Based PlanningJaehyeon Son, Soochan Lee, Gunhee KimICLR 2025
- Vintix II: Decision Pre-Trained Transformer is a Scalable In-Context Reinforcement LearnerAndrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Artyom Grishin et al.ICLR 2026
- Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised PretrainingLicong Lin, Yu Bai, Song MeiICLR 2024 · 74 citations
- Causal Decision Transformer for Recommender Systems via Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Dietmar Jannach, Lina YaoSIGIR 2023 · 33 citations
