ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
Qiang Zhang, Boli Chen, Fanrui Zhang, Ruixue Ding, Shihang Wang, Qiuchen Wang, Yinfeng Huang, Haonan Zhang, Rongxiang Zhu, Xin Li, Houquan zhou, Pengjun Xie
Abstract
Reinforcement learning (RL) has substantially improved the performance of large language model (LLM) agents on tasks with verifiable outcomes, such as mathematics and code generation, but it still struggles on open-ended agent tasks with vast solution spaces (e.g., complex travel planning). Due to the absence of objective ground-truth for these tasks, current RL algorithms largely rely on reward models that assign scalar scores to individual responses. We contend that such pointwise scoring suffers from an inherent discrimination collapse: the reward model struggles to distinguish subtle advantages among different trajectories, resulting in scores within a group being compressed into a narrow range. Consequently, the effective reward signal becomes dominated by noise from the reward model, leading to optimization stagnation. To tackle this issue, we propose ArenaRL, a reinforcement learning paradigm that shifts from pointwise scalar scoring to intra-group relative ranking. ArenaRL introduces a process-aware pairwise evaluation mechanism, employing multi-level rubrics to assign fine-grained relative scores to trajectories. Additionally, we construct an intra-group adversarial arena and devise a tournament-based ranking scheme to obtain stable advantage signals. Empirical results confirm that the built seeded single-elimination scheme achieves nearly equivalent advantage estimation accuracy to full pairwise comparisons with O(N 2 ) complexity, while operating with only O(N) complexity, striking an optimal balance between efficiency and precision. Furthermore, to address the lack of full-cycle benchmarks for open-ended agents, we build Open-Travel and Open-DeepResearch, two high-quality benchmarks featuring a comprehensive pipeline covering supervised fine-tuning, RL training, and multi-dimensional automated evaluation. Extensive experiments show that ArenaRL substantially outperforms standard RL baselines, enabling LLM agents to generate solutions that are more logically rigorous and robust on complex real-world tasks. The code is available at https://github.com/Alibaba-NLP/qqr .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d330050-189d-4f2a-af35-0b759b78f1eeCited by top-tier papers4
- Reward Modeling from Natural Language Human FeedbackZongqi Wang, Rui Wang, Yuchuan Wu, Yiyao Yu et al.ICML 2026 · 6 citations
- Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Non-Verifiable DomainsTejas Krishnan, Sumeet Motwani, Charles London, Suhaas Bhat et al.ICML 2026
- MindFlow: Mind Supernet Powered Thinking Flows for Research Idea InnovationMengdi Liu, Wenjue Chen, Wenyue Chen, Cheng Yang et al.ICML 2026
- GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward ModelJingyu Zhang, Kun Yang, Ming Wen, jiawei zhao et al.ICML 2026
Builds on11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao et al.ICLR 2026 · 146 citations
Related papers
- QuRL: Rubrics As Judge For Open-Ended Question AnsweringXiyu Wei, Qingwei Zong, Xiaoguang Li, Eugene J. Yu et al.ICLR 2026
- From Absolute to Relative: Rethinking Reward Shaping in Group-Based Reinforcement LearningWenzhe Niu, Wei He, Zongxia Xie, Jinpeng Ou et al.ICML 2026 · 1 citation
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for Open-Ended LLM ReasoningYang Zhou, Sunzhu Li, Shunyu Liu, Wenkai Fang et al.ICML 2026 · 44 citations
- Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric RewardsJiajie Zhang, Xin Lv, Ling Feng, Lei Hou et al.ACL 2026 · 8 citations
- REAL: Regression-Aware Reinforcement Learning for LLM-as-a-JudgeYasi Zhang, Tianyu Chen, Mingyuan Zhou, Oscar Leong et al.ICML 2026
