V1: Unifying Generation and Self-Verification for Parallel Reasoners
Harman Singh, Xiuyu Li, Kusha Sareen, Monishwaran Maheswaran, Sijun Tan, Xiaoxia (Shirley) Wu, Junxiong Wang, Alpay Ariyak, Qingyang Wu, Samir Khaki, Rishabh Tiwari, Long (Tony) Lian
摘要
Test-time scaling for complex reasoning tasks shows that leveraging inference-time compute, by methods such as independently sampling and aggregating multiple solutions, results in significantly better task outcomes. However, a critical bottleneck is verification : sampling is only effective if correct solutions can be reliably identified among candidates. While existing approaches typically evaluate candidates independently via scalar scoring, we demonstrate that models are substantially stronger at pairwise verification . Leveraging this insight, we introduce V₁ , a framework that unifies generation and verification through efficient pairwise ranking. V₁ comprises two components: V₁-Infer , an uncertainty-guided algorithm using a tournament-based ranking that dynamically allocates self-verification compute to candidate pairs whose relative correctness is most uncertain; and V₁-PairRL , an RL framework that jointly trains a single model as both generator and pairwise self-verifier, ensuring the verifier adapts to the generator's evolving distribution. On competitive code generation, SWE and terminal-based agents, math reasoning, and medical QA benchmarks, V₁-Infer improves Pass@1 by up to 10% over pointwise verification and outperforms SoTA aggregation-based test-time scaling while being more efficient. Furthermore, V₁-PairRL achieves test-time scaling gains of 3.8–10pp over standard RL and pointwise joint training on code, math, and non-verifiable domains, and improves base Pass@1 by up to 8.7pp over standard RL. Project page & code .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
相关 Paper
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui 等NeurIPS 2025 · 被引用 20 次
- What If We Allocate Test-Time Compute Adaptively?Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Ali Subhan 等ICML 2026 · 被引用 3 次
- ReVeal: Self-Evolving Code Agents via Reliable Self-VerificationYiyang Jin, Kunzhao Xu, Hang Li, Xueting Han 等ICLR 2026 · 被引用 13 次
- Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference OptimizationXinyu Qiu, Heng Jia, Zhengwen Zeng, Shuheng Shen 等CVPR 2026 · 被引用 4 次
- S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement LearningRuotian Ma, Peisong Wang, Cheng Liu, Xingyan Liu 等ACL 2025 · 被引用 13 次
