Scalable In-Context Q-Learning
Jinmei Liu, Fuhong Liu, Zhenhong Sun, Jianye HAO, Huaxiong Li, Bo Wang, Daoyi Dong, Chunlin Chen, Zhi Wang
Abstract
Recent advancements in language models have demonstrated remarkable in-context learning abilities, prompting the exploration of in-context reinforcement learning (ICRL) to extend the promise to decision domains. Due to involving more complex dynamics and temporal correlations, existing ICRL approaches may face challenges in learning from suboptimal trajectories and achieving precise in-context inference. In the paper, we propose Scalable In-Context Q-Learning (S-ICQL), an innovative framework that harnesses dynamic programming and world modeling to steer ICRL toward efficient reward maximization and task generalization, while retaining the scalability and stability of supervised pretraining. We design a prompt-based multi-head transformer architecture that simultaneously predicts optimal policies and in-context value functions using separate heads. We pretrain a generalized world model to capture task-relevant information, enabling the construction of a compact prompt that facilitates fast and precise in-context inference. During training, we perform iterative policy improvement by fitting a state value function to an upper-expectile of the Q-function, and distill the in-context value functions into policy extraction using advantage-weighted regression. Extensive experiments across a range of discrete and continuous environments show consistent performance gains over various types of baselines, especially when learning from suboptimal data. Our code is available at https://github.com/NJU-RL/SICQL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18e5fea1-d6a7-420e-9820-cfe8c6e0579cCited by top-tier papers3
- Mixture-of-Experts Meets In-Context Reinforcement LearningWenhao Wu, Fuhong Liu, Haoru Li, Zican Hu et al.NeurIPS 2025 · 15 citations
- Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language SupervisionShilin Zhang, Zican Hu, Wenhao Wu, Xinyi Xie et al.NeurIPS 2025 · 7 citations
- DADP: Domain Adaptive Diffusion PolicyPengcheng Wang, Qinghang Liu, Haotian Lin, Yiheng Li et al.ICML 2026 · 1 citation
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov et al.ICLR 2024 · 816 citations
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang et al.NeurIPS 2025 · 310 citations
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak et al.NeurIPS 2023 · 170 citations
Related papers
- In-Context Compositional Q-Learning for Offline Reinforcement LearningQiushui Xu, Yuhao Huang, Yushu Jiang, Wenliang Zheng et al.ICLR 2026
- Towards Provable Emergence of In-Context Reinforcement LearningJiuqi Wang, Rohan Chandra, Shangtong ZhangNeurIPS 2025 · 5 citations
- Vintix: Action Model via In-Context Reinforcement LearningAndrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Ilya Zisman et al.ICML 2025
- Reward Is Enough: LLMs Are In-Context Reinforcement LearnersKefan Song, Amir Moeini, Peng Wang, Lei Gong et al.ICLR 2026 · 42 citations
- Distilling Reinforcement Learning Algorithms for In-Context Model-Based PlanningJaehyeon Son, Soochan Lee, Gunhee KimICLR 2025
