CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
Shudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao, Songyang Gao, Chengqi Lyu, Yuzhe Gu, Wenwei Zhang, Derek F. Wong, Songyang Zhang, Kai Chen
Abstract
Answer verification is crucial not only for evaluating large language models (LLMs) by matching their unstructured outputs against standard answers, but also serves as the reward model to guide LLM optimization. Most evaluation frameworks rely on regularized matching or employ general LLMs for answer verification, which demands extensive, repetitive customization for regex rules or evaluation prompts. Two fundamental limitations persist in current methodologies: 1) the absence of comprehensive benchmarks that systematically evaluate verification capabilities across different LLMs; and 2) the nascent stage of verifier development, where existing approaches lack both the robustness to handle complex edge cases and the generalizability across different domains. In this work, we develop CompassVerifier, an accurate and robust lightweight verifier model for evaluation and outcome reward. It demonstrates multi-domain competency spanning math, knowledge, and diverse reasoning tasks, with the capability to process various answer types, including multisubproblems, formulas, and sequence answers, while effectively identifying abnormal/invalid responses. We introduce VerifierBench benchmark comprising model outputs collected from multiple data sources, augmented through manual analysis of meta error patterns to enhance CompassVerifier. We anticipate that Com-passVerifier and VerifierBench will facilitate answer verification, evaluation protocols, and reinforcement learning research. Code and dataset are available at https://github.com/ open-compass/CompassVerifier .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Rethinking Verification for LLM Code Generation: From Generation to TestingZihan Ma, Taolin Zhang, Maosong Cao, Junnan Liu et al.NeurIPS 2025 · 19 citations
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language ModelsYuchen Yan, Jin Jiang, Zhenbang Ren, Yijun Li et al.ICLR 2026 · 18 citations
- Don’t Pass@k: A Bayesian Framework for Large Language Model EvaluationMohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin ChaudharyICLR 2026 · 18 citations
- Variation in Verification: Understanding Verification Dynamics in Large Language ModelsYefan Zhou, Austin Xu, Yilun Zhou, Janvijay Singh et al.ICLR 2026 · 17 citations
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li et al.ACL 2026 · 8 citations
Builds on14
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin et al.ACL 2025 · 209 citations
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang et al.NeurIPS 2025 · 153 citations
- Learn to Reason Efficiently with Adaptive Length-based Reward ShapingWei Liu, Ruochen Zhou, Yiyun Deng, Yuzhen Huang et al.ICLR 2026 · 88 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
Related papers
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across DomainsXuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo et al.AAAI 2026 · 16 citations
- SCI-Verifier: Scientific Verifier with ThinkingShenghe Zheng, Chenyu Huang, Fangchen Yu, Junchi Yao et al.ICLR 2026 · 5 citations
- Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse DomainsYi Su, Dian Yu, Linfeng Song, Juntao Li et al.ACL 2026
- Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And SelectionSadegh Mahdavi, Branislav Kisacanin, Shubham Toshniwal, Wei Du et al.ICML 2026 · 10 citations
- Reinforcing General Reasoning Without VerifiersXiangxin Zhou, Zichen Liu, Anya Sims, Haonan Wang et al.ICLR 2026 · 75 citations
