Beyond the Context Window: Scaling Agentic RL via End-to-end Optimized Context Compression
Miao Lu, Weiwei Sun, Weihua Du, Zhan Ling, Xuesong Yao, Kang Liu, Jiecao Chen
摘要
We study reinforcement learning (RL) finetuning of large language model (LLM) agents for long-horizon multi-turn tool use, where context length quickly becomes a fundamental bottleneck. Existing multi-turn RL pipelines suffer from degraded instruction following, excessive rollout costs, and most importantly, strict context limits. In this work, to address these challenges, we introduce summarization-based context management to training. In specific, it periodically compresses the tool using history by LLM-generated summaries that retain taskrelevant information to keep a compact context while enabling the agent to scale beyond the fixed context window. Building on this formulation, we derive a policy gradient representation that seamlessly enables standard LLM RL infrastructures to optimize both tool-use behaviors as well as summarization strategies in an end-to-end fashion. We instantiate this framework with SUmmarization augmented Policy Optimization (SUPO), an LLM RL algorithm that enables long-horizon training beyond a fixed context limit. Experiments on interactive function calling and searching tasks demonstrate that SUPO significantly improves the success rate while maintaining the same or even lower working context length compared to baselines. We also demonstrate that for complex searching tasks SUPO can further improve the evaluation performance when scaling test-time maximum round of summarization beyond that of training time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao 等NeurIPS 2025 · 被引用 1,138 次
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang 等NeurIPS 2025 · 被引用 387 次
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement LearningSikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie 等ACL 2026 · 被引用 140 次
相关 Paper
- Scaling Long-Horizon Agent via Context FoldingWeiwei Sun, Lu Miao, Zhan Ling, Kang Liu 等ICML 2026 · 被引用 104 次
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RLYifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine 等ICML 2024 · 被引用 163 次
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsYi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan 等ACL 2026 · 被引用 40 次
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search AgentsGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan 等ICLR 2026 · 被引用 37 次
- Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHFZhaolin Gao, Wenhao Zhan, Jonathan Daniel Chang, Gokul Swamy 等ICLR 2025
