Rationale-Aware Answer Verification by Pairwise Self-Evaluation
Akira Kawabata, Saku Sugawara
Abstract
Answer verification identifies correct solutions among candidates generated by large language models (LLMs). Current approaches typically train verifier models by labeling solutions as correct or incorrect based solely on whether the final answer matches the gold answer. However, this approach neglects any flawed rationale in the solution yielding the correct answer, undermining the verifier's ability to distinguish between sound and flawed rationales. We empirically show that in StrategyQA, only 19% of LLM-generated solutions with correct answers have valid rationales. Furthermore, we demonstrate that training a verifier on valid rationales significantly improves its ability to distinguish valid and flawed rationales. To make a better verifier without extra human supervision, we introduce REPS (Rationale Enhancement through Pairwise Selection), a method for selecting valid rationales from candidates by iteratively applying pairwise self-evaluation using the same LLM that generates the solutions. Verifiers trained on solutions selected by REPS outperform those trained using conventional training methods on three reasoning benchmarks (ARC-Challenge, DROP, and StrategyQA). Our results suggest that training reliable verifiers requires ensuring the validity of rationales in addition to the correctness of the final answers, which would be critical for models assisting humans in solving complex reasoning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Provable Scaling Laws for the Test-Time Compute of Large Language ModelsYanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding et al.NeurIPS 2025 · 16 citations
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught ReasonersReiss Koh, Wonbeen Oh, Jaein Jang, Minhyung Lee et al.NeurIPS 2025 · 8 citations
- SkillVerse : Assessing and Enhancing LLMs with Tree EvaluationYufei Tian, Jiao Sun, Nanyun Peng, Zizhao ZhangACL 2025 · 1 citation
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome RewardShudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao et al.EMNLP 2025
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale CalibrationYuanchen Wu, Ke Yan, Shouhong Ding, Ziyin Zhou et al.ICML 2025
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable RewardsXiaoyuan Liu, Tian Liang, Zhiwei He, Jiahao Xu et al.NeurIPS 2025 · 46 citations
- Generative Verifiers: Reward Modeling as Next-Token PredictionLunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi et al.ICLR 2025
- Learning to Self-Verify Makes Language Models Better ReasonersYuxin Chen, Yu Wang, Yi Zhang, Ziang Ye et al.ICML 2026 · 12 citations
- ORACLE: Optimizing Reasoning Abilities of Large Language Models via Constraint-Led Synthetic Data ElicitationZhuojie Yang, Wentao Wan, Keze WangAAAI 2026
