C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
Akira Kawabata, Saku Sugawara
摘要
Rubric-augmented verification guides reward models with explicit evaluation criteria, yielding more reliable judgments than single-model verification. However, most existing methods require costly rubric annotations, limiting scalability. Moreover, we find that rubric generation is vulnerable to a failure of cooperation; lowquality rubrics actively mislead reward models rather than help. Inspired by the principle of cooperative communication, we propose Cooperative yet Critical reward modeling (C2), a framework that significantly improves reward model judgments by having the reward model critically collaborate with a rubric generator trained solely from binary preferences. In C2, we synthesize helpful and misleading rubric pairs by measuring how each rubric shifts the reward model toward or away from the correct preference. Using these contrastive pairs, we train a cooperative rubric generator to propose helpful rubrics, and a critical verifier to assess rubric validity before making its judgment, following only rubrics it deems helpful at inference time. C2 outperforms reasoning reward models trained on the same binary preferences, with gains of up to 6.5 points on RM-Bench and 6.0 points length-controlled win rate on Al-pacaEval 2.0. Without external rubric annotations, C2 enables an 8B reward model to match performance achieved with rubrics from a 4× larger model. Overall, our work demonstrates that eliciting deliberate cooperation in rubricaugmented verification makes reward models more trustworthy in a scalable way. 1 * Work done while at The Asahi Shimbun Company. 1 Our code is available at https://github.com/ asahi-research/C2 . Verifier (a.k.a. Reward Model) Prompt Respond politely to the user's complaint and offer a solution. User: "I've been waiting for my refund ...
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsAnisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath 等ICLR 2026 · 被引用 340 次
- ULTRAFEEDBACK: Boosting Language Models with Scaled AI FeedbackGanqu Cui, Lifan Yuan, Ning Ding, Guanming Yao 等ICML 2024 · 被引用 286 次
相关 Paper
- RubricBench: Aligning Model-Generated Rubrics with Human StandardsJunyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu 等ACL 2026 · 被引用 7 次
- CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward ModelingDengcan Liu, Fengkai Yang, Xiaohan Wang, Shurui Yan 等KDD 2026 · 被引用 12 次
- Robust Reward Modeling via Causal RubricsPragya Srivastava, Harman Singh, Rahul Madhavan, Gandharv Patil 等ICLR 2026 · 被引用 20 次
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation TasksMian Wu, Gavin Zhang, Sewon Min, Sergey Levine 等ICLR 2026 · 被引用 15 次
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentTianci Liu, Ran Xu, Tony Yu, Ilgee Hong 等ACL 2026 · 被引用 75 次
