When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
Shayan Kiyani, Sima Noorani, George Pappas, Hamed Hassani
摘要
Reasoning with LLMs increasingly unfolds inside a broader verification loop. Internally, systems use cheap checks, such as self-consistency or proxy rewards, which we call weak verification . Externally, users inspect outputs and steer the model through feedback until results are trustworthy, which we call strong verification . These signals differ sharply in cost and reliability: strong verification can establish trust but is resource-intensive, while weak verification is fast and scalable but noisy and imperfect. We formalize this tension through weak-strong verification policies , which decide when to accept or reject based on weak verification and when to defer to strong verification. We introduce metrics capturing incorrect acceptance, incorrect rejection, and strong-verification frequency. Over population, we show that optimal policies admit a two-threshold structure and that calibration and sharpness govern the value of weak verifiers. Building on this, we develop an online algorithm that provably controls acceptance and rejection errors without assumptions on the query stream, the language model, or the weak verifier. Experiments on mathematical reasoning and sequential decision-making demonstrate that our algorithm achieves reliability comparable to exhaustive strong verification while significantly reducing verification cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Adaptive Conformal Inference Under Distribution ShiftIsaac Gibbs, Emmanuel J. CandèsNeurIPS 2021 · 被引用 665 次
- Self-Evaluation Guided Beam Search for ReasoningYuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao 等NeurIPS 2023 · 被引用 316 次
相关 Paper
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable RewardsXiaoyuan Liu, Tian Liang, Zhiwei He, Jiahao Xu 等NeurIPS 2025 · 被引用 46 次
- LaSeR: Reinforcement Learning with Last-Token Self-RewardingWenkai Yang, Weijie Liu, Ruobing Xie, Yiju Guo 等ICLR 2026 · 被引用 12 次
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative VerifierJianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen 等ACL 2026 · 被引用 3 次
- Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable RewardsReinhard Heckel, Mahdi Soltanolkotabi, Christos ThrampoulidisICML 2026 · 被引用 2 次
- Dyve: Thinking Fast and Slow for Dynamic Process VerificationJianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen 等EMNLP 2025
