Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across Modalities
Chi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin, Jiaming Ji, Juntao Dai, Wei Xue, Sirui Han, Yike Guo
摘要
The rise of Omni-modality Large Language Models (OLLMs) capable of jointly processing text, audio, and visual inputs marks a major step toward general intelligence. Ensuring their alignment with human preferences requires effective Omni-modality Reward Models (ORMs), which serve as surrogates for human judgment to guide OLLMs behavior. However, ORMs evaluation remains under-developed in the previous literature. Existing benchmarks are largely text-centric or limited to bimodal tasks, restricting comprehensive assessment for ORMs. To bridge this gap, we introduce Omni-RewardBench , the first benchmark for comprehensive evaluation of ORMs across modalities. In short, our contributions are threefold: (1) a hybrid automatic-annotation and human-verification pipeline to construct high-quality evaluation data; (2) extensive experiments on 20+ models, including inherently omni-modal and modality-bridged systems. Our experimental results demonstrate that current OLLMs fall short as reward models, revealing several common failure modes such as perception failure , modality dominance failure , and cross-modal fusion failure . and (3) strong correlations between Omni-RewardBench scores and downstream performance (IID r = 0.94, OOD r = 0.72), validating its reliability as a predictor of real-world capability and alignment quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin 等ACL 2025 · 被引用 209 次
- PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward ModelsMingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou 等ACL 2025 · 被引用 85 次
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and ActionJiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang 等CVPR 2024 · 被引用 53 次
相关 Paper
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form PreferencesZhuoran Jin, Hongbang Yuan, Kejian Zhu, Jiachun Li 等ICLR 2026 · 被引用 11 次
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu 等ICLR 2026 · 被引用 4 次
- AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMsYaoting Wang, Ziyi Zhang, Wenming Tu, Shaoxuan Xu 等ICML 2026
- OmnixR: Evaluating Omni-modality Language Models on Reasoning across ModalitiesLichang Chen, Hexiang Hu, Mingda Zhang, Yiwen Chen 等ICLR 2025
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning EvaluationJianghan Chao, Jianzhang Gao, Wenhui Tan, Yuchong Sun 等ICLR 2026 · 被引用 16 次
