Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning
Jaehyeon Son, Soochan Lee, Gunhee Kim
Abstract
Recent studies have shown that Transformers can perform in-context reinforcement learning (RL) by imitating existing RL algorithms, enabling sample-efficient adaptation to unseen tasks without parameter updates. However, these models also inherit the suboptimal behaviors of the RL algorithms they imitate. This issue primarily arises due to the gradual update rule employed by those algorithms. Modelbased planning offers a promising solution to this limitation by allowing the models to simulate potential outcomes before taking action, providing an additional mechanism to deviate from the suboptimal behavior. Rather than learning a separate dynamics model, we propose Distillation for In-Context Planning (DICP), an in-context model-based RL framework where Transformers simultaneously learn environment dynamics and improve policy in-context. We evaluate DICP across a range of discrete and continuous environments, including Darkroom variants and Meta-World. Our results show that DICP achieves state-of-the-art performance while requiring significantly fewer environment interactions than baselines, which include both model-free counterparts and existing meta-RL methods. The code is available at https://github.com/jaehyeon-son/dicp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a405704-0088-46a5-bb65-da955fc7455cCited by top-tier papers4
- Mixture-of-Experts Meets In-Context Reinforcement LearningWenhao Wu, Fuhong Liu, Haoru Li, Zican Hu et al.NeurIPS 2025 · 15 citations
- Scalable In-Context Q-LearningJinmei Liu, Fuhong Liu, Zhenhong Sun, Jianye HAO et al.ICLR 2026 · 8 citations
- Safe In-Context Reinforcement LearningAmir Moeini, Minjae Kwon, Alper Bozkurt, Yuichi Motai et al.ICML 2026 · 4 citations
- Vintix II: Decision Pre-Trained Transformer is a Scalable In-Context Reinforcement LearnerAndrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Artyom Grishin et al.ICLR 2026
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
Related papers
- Efficient Cross-Episode Meta-RLGresa Shala, André Biedenkapp, Pierre Krack, Florian Walter et al.ICLR 2025
- Towards Provable Emergence of In-Context Reinforcement LearningJiuqi Wang, Rohan Chandra, Shangtong ZhangNeurIPS 2025 · 5 citations
- Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised PretrainingLicong Lin, Yu Bai, Song MeiICLR 2024 · 74 citations
- Transformers are Meta-Reinforcement LearnersLuckeciano C. MeloICML 2022 · 66 citations
- In-context Reinforcement Learning with Algorithm DistillationMichael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto et al.ICLR 2023 · 10 citations
