Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization
Xingyuan Hua, Sheng Yue, Ju Ren
摘要
Recent advancements in agentic test-time scaling allow models to gather environmental feedback before committing to final actions. A key limitation of existing methods is that they typically employ undifferentiated exploration strategies, lacking the ability to adaptively distinguish when exploration is truly required. In this paper, we propose an exploration-aware reinforcement learning framework that enables LLM agents to adaptively explore only when uncertainty is high. Our method introduces a fine-grained reward function via variational inference that explicitly evaluates exploratory actions by estimating their potential to improve future decision-making, together with an exploration-aware grouping mechanism that separates exploratory actions from task-completion actions during optimization. By targeting informational gaps, this design allows agents to explore selectively and transition to execution as soon as the task context is clear. Empirically, we demonstrate that our approach achieves consistent improvements across a range of challenging text-based and GUI-based agent benchmarks. Code is available at https://github. com/HansenHua/EAPO-ICML26 and models are available at https://huggingface. co/hansenhua/EAPO-ICML26 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
相关 Paper
- Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied AgentsSeohui Bae, Jeonghye Kim, Youngchul Sung, Woohyung LimCVPR 2026 · 被引用 1 次
- SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action ModelsHyeonbeom Choi, Daechul Ahn, Youhan Lee, Taewook Kang 等ICML 2026 · 被引用 2 次
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao 等ICLR 2026 · 被引用 146 次
- Timely Machine: Awareness of Time Makes Test-Time Scaling AgenticYichuan Ma, Linyang Li, Yongkang Chen, Peiji Li 等ACL 2026 · 被引用 3 次
- On-the-Fly VLA Adaptation via Test-Time Reinforcement LearningChangyu Liu, Yiyang Liu, Taowen Wang, Qiao Zhuang 等ACL 2026 · 被引用 7 次
