Beyond the Context Window: Scaling Agentic RL via End-to-end Optimized Context Compression
Miao Lu, Weiwei Sun, Weihua Du, Zhan Ling, Xuesong Yao, Kang Liu, Jiecao Chen
Abstract
We study reinforcement learning (RL) finetuning of large language model (LLM) agents for long-horizon multi-turn tool use, where context length quickly becomes a fundamental bottleneck. Existing multi-turn RL pipelines suffer from degraded instruction following, excessive rollout costs, and most importantly, strict context limits. In this work, to address these challenges, we introduce summarization-based context management to training. In specific, it periodically compresses the tool using history by LLM-generated summaries that retain taskrelevant information to keep a compact context while enabling the agent to scale beyond the fixed context window. Building on this formulation, we derive a policy gradient representation that seamlessly enables standard LLM RL infrastructures to optimize both tool-use behaviors as well as summarization strategies in an end-to-end fashion. We instantiate this framework with SUmmarization augmented Policy Optimization (SUPO), an LLM RL algorithm that enables long-horizon training beyond a fixed context limit. Experiments on interactive function calling and searching tasks demonstrate that SUPO significantly improves the success rate while maintaining the same or even lower working context length compared to baselines. We also demonstrate that for complex searching tasks SUPO can further improve the evaluation performance when scaling test-time maximum round of summarization beyond that of training time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 263c2224-1787-45a9-88e3-58fef034fd40Builds on13
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao et al.NeurIPS 2025 · 1,138 citations
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye et al.AAAI 2024 · 394 citations
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang et al.NeurIPS 2025 · 387 citations
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement LearningSikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie et al.ACL 2026 · 140 citations
Related papers
- Scaling Long-Horizon Agent via Context FoldingWeiwei Sun, Lu Miao, Zhan Ling, Kang Liu et al.ICML 2026 · 104 citations
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RLYifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine et al.ICML 2024 · 163 citations
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsYi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan et al.ACL 2026 · 40 citations
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search AgentsGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan et al.ICLR 2026 · 37 citations
- Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHFZhaolin Gao, Wenhao Zhan, Jonathan Daniel Chang, Gokul Swamy et al.ICLR 2025
