Long-form RewardBench: Evaluating Reward Models for Long-form Generation
Hui Huang, Yancheng He, Wei Liu, Muyun Yang, Jiaheng Liu, Kehai Chen, Bing Xu, Conghui Zhu, Hailong Cao, Tiejun Zhao
摘要
The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for longform generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Longform Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifiers exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
- Self-Alignment with Instruction BacktranslationXian Li, Ping Yu, Chunting Zhou, Timo Schick 等ICLR 2024 · 被引用 174 次
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu 等ACL 2024 · 被引用 18 次
- Think-J: Learning to Think for Generative LLM-as-a-JudgeHui Huang, Yancheng He, Hongli Zhou, Rui Zhang 等AAAI 2026 · 被引用 12 次
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement LearningYuhao Wu, Yushi Bai, Zhiqiang Hu, Roy Ka-Wei Lee 等ICLR 2026 · 被引用 9 次
相关 Paper
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMsYuhao Wu, Ming Shan Hee, Zhiqiang Hu, Roy Ka-Wei LeeICLR 2025
- M-RewardBench: Evaluating Reward Models in Multilingual SettingsSrishti Gureja, Lester James Validad Miranda, Shayekh Bin Islam, Rishabh Maheshwary 等ACL 2025
- Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward ModelingJiaxuan Wang, Yulan Hu, Wenjin Yang, Zheng Pan 等ACL 2026 · 被引用 1 次
- RMB: Comprehensively benchmarking reward models in LLM alignmentEnyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi 等ICLR 2025
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward ModelsChenglong Wang, Yifu Huo, Yang Gan, Yongyu Mu 等AAAI 2026 · 被引用 1 次
