SOAR: Supervision from Observation for Agentic Reinforcement Learning
Meng Li, Lei Li, Xiting Wang, Yi Yuan, Zheng Wei, Jiang Bian, Zang Li
摘要
Agentic reinforcement learning enables large language models to solve long-horizon tasks by interacting with the environment and internalizing tool-use behavior into their reasoning. Prior work assigns supervision primarily based on outcome rewards or external reward models, but largely ignores environment observations, a critical source of learning. Consequently, agents may identify successful actions without understanding how the environment responds, producing suboptimal policies. To address this, we propose SOAR (Supervision from Observation for Agentic Reinforcement Learning), which assigns positive advantages to observation tokens proportional to the negative entropy of preceding actions. This encourages the agent to learn from outcomes of confident actions, grounding policy updates in environment dynamics and improving anticipation of tool-call consequences. Empirical results across three domains and 13 benchmarks show that SOAR consistently improves performance, yielding gains of up to 7.0% on general reasoning tasks and 16.9% on deep research tasks, while reducing erroneous and inefficient tool usage. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
相关 Paper
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao 等ICLR 2026 · 被引用 146 次
- Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model AgentsZeping Li, Hongru Wang, Yiwen Zhao, Guanhua Chen 等ACL 2026
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise ReasoningYihe Deng, I-Hung Hsu, Jun Yan, Zifeng Wang 等ICLR 2026 · 被引用 11 次
- ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language ModelsZihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai 等ICLR 2026 · 被引用 10 次
- Toward Generalized Web Agent Training: A Deep Dive into Entropy-Balanced Reinforcement LearningGuanting Dong, Licheng Bao, Zhongyuan Wang, Kangzhi Zhao 等WWW 2026 · 被引用 2 次
