SOAR: Supervision from Observation for Agentic Reinforcement Learning
Meng Li, Lei Li, Xiting Wang, Yi Yuan, Zheng Wei, Jiang Bian, Zang Li
Abstract
Agentic reinforcement learning enables large language models to solve long-horizon tasks by interacting with the environment and internalizing tool-use behavior into their reasoning. Prior work assigns supervision primarily based on outcome rewards or external reward models, but largely ignores environment observations, a critical source of learning. Consequently, agents may identify successful actions without understanding how the environment responds, producing suboptimal policies. To address this, we propose SOAR (Supervision from Observation for Agentic Reinforcement Learning), which assigns positive advantages to observation tokens proportional to the negative entropy of preceding actions. This encourages the agent to learn from outcomes of confident actions, grounding policy updates in environment dynamics and improving anticipation of tool-call consequences. Empirical results across three domains and 13 benchmarks show that SOAR consistently improves performance, yielding gains of up to 7.0% on general reasoning tasks and 16.9% on deep research tasks, while reducing erroneous and inefficient tool usage. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
Related papers
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao et al.ICLR 2026 · 146 citations
- Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model AgentsZeping Li, Hongru Wang, Yiwen Zhao, Guanhua Chen et al.ACL 2026
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise ReasoningYihe Deng, I-Hung Hsu, Jun Yan, Zifeng Wang et al.ICLR 2026 · 11 citations
- ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language ModelsZihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai et al.ICLR 2026 · 10 citations
- Toward Generalized Web Agent Training: A Deep Dive into Entropy-Balanced Reinforcement LearningGuanting Dong, Licheng Bao, Zhongyuan Wang, Kangzhi Zhao et al.WWW 2026 · 2 citations
