Rationale-Aware Answer Verification by Pairwise Self-Evaluation
Akira Kawabata, Saku Sugawara
摘要
Answer verification identifies correct solutions among candidates generated by large language models (LLMs). Current approaches typically train verifier models by labeling solutions as correct or incorrect based solely on whether the final answer matches the gold answer. However, this approach neglects any flawed rationale in the solution yielding the correct answer, undermining the verifier's ability to distinguish between sound and flawed rationales. We empirically show that in StrategyQA, only 19% of LLM-generated solutions with correct answers have valid rationales. Furthermore, we demonstrate that training a verifier on valid rationales significantly improves its ability to distinguish valid and flawed rationales. To make a better verifier without extra human supervision, we introduce REPS (Rationale Enhancement through Pairwise Selection), a method for selecting valid rationales from candidates by iteratively applying pairwise self-evaluation using the same LLM that generates the solutions. Verifiers trained on solutions selected by REPS outperform those trained using conventional training methods on three reasoning benchmarks (ARC-Challenge, DROP, and StrategyQA). Our results suggest that training reliable verifiers requires ensuring the validity of rationales in addition to the correctness of the final answers, which would be critical for models assisting humans in solving complex reasoning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- Provable Scaling Laws for the Test-Time Compute of Large Language ModelsYanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding 等NeurIPS 2025 · 被引用 16 次
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught ReasonersReiss Koh, Wonbeen Oh, Jaein Jang, Minhyung Lee 等NeurIPS 2025 · 被引用 8 次
- SkillVerse : Assessing and Enhancing LLMs with Tree EvaluationYufei Tian, Jiao Sun, Nanyun Peng, Zizhao ZhangACL 2025 · 被引用 1 次
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome RewardShudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao 等EMNLP 2025
它引用的顶会 Paper21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
相关 Paper
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale CalibrationYuanchen Wu, Ke Yan, Shouhong Ding, Ziyin Zhou 等ICML 2025
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable RewardsXiaoyuan Liu, Tian Liang, Zhiwei He, Jiahao Xu 等NeurIPS 2025 · 被引用 46 次
- Generative Verifiers: Reward Modeling as Next-Token PredictionLunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi 等ICLR 2025
- Learning to Self-Verify Makes Language Models Better ReasonersYuxin Chen, Yu Wang, Yi Zhang, Ziang Ye 等ICML 2026 · 被引用 12 次
- ORACLE: Optimizing Reasoning Abilities of Large Language Models via Constraint-Led Synthetic Data ElicitationZhuojie Yang, Wentao Wan, Keze WangAAAI 2026
