SCI-Verifier: Scientific Verifier with Thinking
Shenghe Zheng, Chenyu Huang, Fangchen Yu, Junchi Yao, Jingqi Ye, Tao Chen, Yun Luo, Ning Ding, Lei Bai, Ganqu Cui, Peng Ye
摘要
As large language models (LLMs) are increasingly applied to scientific reasoning, the complexity of answer formats and the diversity of equivalent expressions make answer verification a critical yet challenging task. Existing verification studies in scientific domains suffer from two major limitations: (a) the absence of systematic evaluation standards and insufficient disciplinary coverage, which hinders their comprehensive assessment; and (b) heavy reliance on cumbersome rule design or prompt engineering, which reduces their effectiveness in complex reasoning scenarios or limits their cross-disciplinary generalization. To address these challenges, we propose solutions at both the data and model levels. On the data side, we construct SCI-VerifyBench, a cross-disciplinary benchmark covering mathematics, physics, biology, chemistry, and general scientific QA. The benchmark is built from real LLM responses and enhanced with domain-specific equivalence transformations that generate challenging and realistic data. Model-based and expert annotations ensure both quality and diversity, enabling rigorous evaluation of verification ability. On the model side, we emphasize the importance of reasoning for verification and introduce SCI-Verifier, a unified reasoning-augmented verifier for scientific domains. Through post-training, SCI-Verifier demonstrates strong logical reasoning and equivalence judgment capabilities while maintaining concise and stable outputs. Together, SCI-VerifyBench and SCI-Verifier provide a principled framework for scientific verification, offering both systematic evaluation and practical pathways to enhance the reliability and applicability of LLMs in scientific domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language ModelsHongji Li, Manjiang Yu, Junchi Yao, PRIYANKA SINGH 等CVPR 2026 · 被引用 3 次
- PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and EngineeringXiangfeng Wang, Hangyu Guo, Yanlin Lai, Mitt Huang 等ACL 2026 · 被引用 1 次
- FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision–Language ModelsChenyu Huang, Peng Ye, Xudong Tan, Jinhan Mu 等ICML 2026
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang 等NeurIPS 2025 · 被引用 153 次
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement LearningChenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li 等ICLR 2026 · 被引用 74 次
- Making Language Models Better Reasoners with Step-Aware VerifierYifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu 等ACL 2023 · 被引用 52 次
相关 Paper
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across DomainsXuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo 等AAAI 2026 · 被引用 16 次
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome RewardShudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao 等EMNLP 2025
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu 等ICML 2024 · 被引用 220 次
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code GenerationQiaosheng Chen, Yang Liu, Lei Li, Kai Chen 等ICML 2026 · 被引用 1 次
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language ModelsJiacheng Ruan, Dan Jiang, Xian Gao, Ting Liu 等AAAI 2026 · 被引用 3 次
