Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Yunzhong He, Bing Liu, Sean M. Hendryx
摘要
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it to real-world reasoning tasks is challenging, as evaluation depends on nuanced, multi-criteria judgments rather than binary correctness. Instance-specific rubrics have recently been used in evaluation benchmarks to capture such judgments, but their potential as reward signals for on-policy post-training remains underexplored. We introduce , an on-policy reinforcement learning method that extends RLVR beyond verifiable domains by using rubric-based feedback. Across both medical and science domains, we evaluate multiple strategies for aggregating rubric feedback into rewards. The best RaR variant achieves relative improvements of up to 31% on HealthBench and 7% on GPQA-Diamond over popular LLM-as-judge baselines that rely on direct Likert-based rewards. These results demonstrate that RaR-trained policies adapt well to diverse evaluation formats, performing strongly on both rubric-based and multiple-choice tasks. Moreover, we find that using rubrics as structured reward signals yields better alignment for smaller judges and reduces performance variance across judge scales.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin 等ICLR 2026 · 被引用 147 次
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic WeightingWenhao Zhang, Yuexiang Xie, Yuchang Sun, Yanxi Chen 等ICLR 2026 · 被引用 100 次
- Reinforcement Learning with Evolving Rubrics for Deep ResearchRulin Shao, Akari Asai, Shannon Shen, Hamish Ivison 等ICML 2026 · 被引用 78 次
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentTianci Liu, Ran Xu, Tony Yu, Ilgee Hong 等ACL 2026 · 被引用 75 次
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 被引用 58 次
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine 等ICLR 2026 · 被引用 218 次
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang 等NeurIPS 2025 · 被引用 153 次
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin 等ICLR 2026 · 被引用 147 次
相关 Paper
- Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Non-Verifiable DomainsTejas Krishnan, Sumeet Motwani, Charles London, Suhaas Bhat 等ICML 2026
- Online Rubrics Elicitation from Pairwise ComparisonsMohammadHossein Rezaei, Robert Vacareanu, Zihao Wang, Clinton Wang 等ICML 2026 · 被引用 41 次
- Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain ReasoningBaolong Bi, Shenghua Liu, Yiwei Wang, Siqian Tong 等ICML 2026 · 被引用 19 次
- QuRL: Rubrics As Judge For Open-Ended Question AnsweringXiyu Wei, Qingwei Zong, Xiaoguang Li, Eugene J. Yu 等ICLR 2026
- RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine GenerationSunzhu Li, Jiale Zhao, Huimin Ren, Zhenlin Wei 等ACL 2026 · 被引用 22 次
