Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
yibo li, Zijie Lin, Ailin Deng, Xuan (Billy) Zhang, Yufei He, Shuo Ji, Tri Cao, Bryan Hooi
Abstract
While Large Language Model (LLM) agents excel at general tasks, they inherently struggle with continual adaptation due to the frozen weights after deployment. Conventional reinforcement learning (RL) offers a solution but incurs prohibitive computational costs and the risk of catastrophic forgetting. We introduce Just-In-Time Reinforcement Learning (JitRL), a training-free framework that enables test-time policy optimization without any gradient updates. JitRL maintains a dynamic, non-parametric memory of experiences and retrieves relevant trajectories to estimate action advantages on-the-fly. These estimates are then used to directly modulate the LLM's output logits. We theoretically prove that this additive update rule is the exact closed-form solution to the KL-constrained policy optimization objective. Extensive experiments on WebArena and Jericho demonstrate that JitRL establishes a new state-of-the-art among training-free methods. Crucially, JitRL outperforms the performance of computationally expensive fine-tuning methods (e.g., WebRL) while reducing monetary costs by over 30 times, offering a scalable path for continual learning agents. The code is available at https://anonymous.4open.science/r/JitRL-D485.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aaca1bc6-876a-4775-bfd2-58ff50236c9eCited by top-tier papers1
Ask how each one uses itBuilds on10
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao et al.NeurIPS 2025 · 1,138 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
Related papers
- TMS: Trajectory-Mixed Supervision for On-Policy Self DistillationRana Khan, Zijie Liu, Zhen Tan, Charles Fleming et al.ICML 2026
- Test-Time Learning for Large Language ModelsJinwu Hu, Zitian Zhang, Guohao Chen, Xutao Wen et al.ICML 2025
- Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy OptimizationWengao Ye, Yan Liang, Lianlei ShanICLR 2026 · 12 citations
- In-Place Test-Time TrainingGuhao Feng, Shengjie Luo, Kai Hua, Ge Zhang et al.ICLR 2026 · 15 citations
- Accelerating RL for LLM Reasoning with Optimal Advantage RegressionKianté Brantley, Mingyu Chen, Zhaolin Gao, Jason D. Lee et al.NeurIPS 2025 · 31 citations
