Online Rubrics Elicitation from Pairwise Comparisons
MohammadHossein Rezaei, Robert Vacareanu, Zihao Wang, Clinton Wang, Bing Liu, Yunzhong He, Afra Feyza Akyürek
摘要
Rubrics provide a flexible way to train LLMs on open-ended long-form answers where verifiable rewards are not applicable and human preferences provide coarse signals. Prior work shows that reinforcement learning with rubric-based rewards leads to consistent gains in LLM post-training. Most existing approaches rely on rubrics that remain static over the course of training. Such static rubrics, however, are vulnerable to reward-hacking type behaviors and fail to capture emergent desiderata that arise during training. We introduce Online Rubrics Elicitation (OnlineRubrics), a method that dynamically curates evaluation criteria in an online manner through pairwise comparisons of responses from current and reference policies. This online process enables continuous identification and mitigation of errors as training proceeds. Empirically, this approach yields consistent improvements of up to 8% over training exclusively with static rubrics across AlpacaEval, GPQA, ArenaHard as well as the validation sets of expert questions and rubrics. We qualitatively analyze the elicited criteria and identify prominent themes such as transparency, practicality, organization, and reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Reinforcement Learning with Evolving Rubrics for Deep ResearchRulin Shao, Akari Asai, Shannon Shen, Hamish Ivison 等ICML 2026 · 被引用 78 次
- PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional ReasoningAfra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang, Vipul Gupta 等ACL 2026 · 被引用 18 次
- Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric RewardsJiajie Zhang, Xin Lv, Ling Feng, Lei Hou 等ACL 2026 · 被引用 8 次
- PerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionYana Wei, Hongbo Peng, Yanlin Lai, Liang Zhao 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsAnisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath 等ICLR 2026 · 被引用 340 次
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye 等ICLR 2026 · 被引用 279 次
相关 Paper
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentTianci Liu, Ran Xu, Tony Yu, Ilgee Hong 等ACL 2026 · 被引用 75 次
- QuRL: Rubrics As Judge For Open-Ended Question AnsweringXiyu Wei, Qingwei Zong, Xiaoguang Li, Eugene J. Yu 等ICLR 2026
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation TasksMian Wu, Gavin Zhang, Sewon Min, Sergey Levine 等ICLR 2026 · 被引用 15 次
- SERL: Self-Examining Reinforcement Learning on Open-DomainWeixuan Ou, Yanzhao Zheng, Shuoshuo Sun, Wei Zhang 等AAAI 2026 · 被引用 1 次
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-TrainingJunkai Zhang, Zihao Wang, Lin Gui, Swarnashree Mysore Sathyendra 等ICLR 2026 · 被引用 48 次
