ARIA: Training Language Agents with Intention-driven Reward Aggregation
Ruihan Yang, Yikai Zhang, Aili Chen, Xintao Wang, Jiangjie Chen, Siyu Yuan, Deqing Yang, Yanghua Xiao
Abstract
Large language models (LLMs) have enabled agents to perform complex reasoning and decision-making through free-form language interactions. However, in open-ended language action environments (e.g., negotiation or question-asking games), the action space can be formulated as a joint distribution over tokens, resulting in an exponentially large action space. Sampling actions in such a space can lead to extreme reward sparsity, which brings large reward variance, hindering effective reinforcement learning (RL). To address this, we propose ARIA, a method that Aggregates Rewards in Intention space to enable efficient and effective language Agents training. ARIA aims to project natural language actions from the high-dimensional joint token distribution space into a low-dimensional intention space, where semantically similar actions are clustered and assigned shared rewards. This intention-aware reward aggregation reduces reward variance by densifying reward signals, fostering better policy optimization. Extensive experiments demonstrate that ARIA not only significantly reduces policy gradient variance, but also delivers substantial performance gains of an average of 9.95% across four downstream tasks, consistently outperforming offline and online RL baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f952035-858d-4f08-9393-39f04ee71f77Cited by top-tier papers2
- Intrinsic Credit Assignment for Long Horizon InteractionIlze Amanda Auzina, Joschka Strüber, Sergio Hernández-Gutiérrez, Shashwat Goel et al.ICML 2026 · 6 citations
- Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM AgentsRuihan Yang, Fanghua Ye, Xiang Wei, Ruoqing Zhao et al.ICML 2026 · 2 citations
Builds on12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari et al.ICLR 2024 · 359 citations
Related papers
- From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent TrainingZishang Jiang, tingyun li, Jinyi Han, Xinyi Wang et al.ICML 2026
- Retroformer: Retrospective Large Language Agents with Policy Gradient OptimizationWeiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu et al.ICLR 2024 · 124 citations
- Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsYifan Zhou, Sachin Grover, Mohamed El Mistiri, Kamalesh Kalirathinam et al.NeurIPS 2025 · 3 citations
- RF-Agent: Automated Reward Function Design via Language Agent Tree SearchNing Gao, Xiuhui Zhang, Xingyu Jiang, Mukang You et al.NeurIPS 2025 · 8 citations
- Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy OptimizationZelai Xu, Wanjun Gu, Chao Yu, Yi Wu et al.ICML 2025
