ExVerus: Verus Proof Repair via Counterexample Reasoning
Jun Yang, Yuechun Sun, Yi Wu, Rodrigo Caridad, Yongwei Yuan, Jianan Yao, Shan Lu, Kexin Pei
Abstract
Large Language Models (LLMs) have shown promising results in automating formal verification. However, existing approaches often treat the proof generation as a static, end-to-end prediction, relying on limited verifier feedback and lacking access to concrete instances of proof failure, i.e., counterexamples, to characterize the discrepancies between the intended behavior specified in the proof and the concrete executions of the code that can violate it. We present EXVERUS, a new framework that enables LLMs to generate and repair Verus proofs with actionable guidance based on the behavioral feedback using counterexamples. When a proof fails, EXVERUS automatically generates counterexamples, and then guides the LLM to learn from counterexamples and block them, incrementally fixing the verification failures. Our evaluation shows that EXVERUS substantially outperforms the state-of-the-art LLMbased proof generator in proof success rate, robustness, cost, and inference efficiency, across a variety of model families, agentic design, error types, and benchmarks with diverse difficulties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba7e23c0-bae3-44c9-baee-03d10c045245Builds on13
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityParshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton et al.NeurIPS 2025 · 507 citations
- Baldur: Whole-Proof Generation and Repair with Large Language ModelsEmily First, Markus N. Rabe, Talia Ringer, Yuriy BrunFSE 2023 · 89 citations
- Verus: Verifying Rust Programs using Linear Ghost TypesAndrea Lattuada, Travis Hance, Chanhee Cho, Matthias Brun et al.OOPSLA 2023 · 86 citations
- Anvil: Verifying Liveness of Cluster Management ControllersXudong Sun, Wenjie Ma, Jiawei Tyler Gu, Zicheng Ma et al.OSDI 2024 · 50 citations
- Towards AI-Assisted Synthesis of Verified Dafny MethodsMd Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James NobleFSE 2024 · 26 citations
Related papers
- AutoVerus: Automated Proof Generation for Rust CodeChenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Jianan Yao et al.OOPSLA 2025 · 11 citations
- PAT-Agent: Autoformalization for Model CheckingXinyue Zuo, Yifan Zhang, Hongshu Wang, Yufan Cai et al.ASE 2025 · 1 citation
- AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and TreefinementPranjal Aggarwal, Bryan Parno, Sean WelleckICML 2025
- Propose, Solve, Verify: Self-Play Through Formal VerificationAlex Wilf, Pranjal Aggarwal, Bryan Parno, Daniel Fried et al.ICML 2026 · 6 citations
- A Tale of 1001 LoC: Potential Runtime Error-Guided Specification Synthesis for Verifying Large-Scale ProgramsZhongyi Wang, Tengjie Lin, Mingshuai Chen, Haokun Li et al.OOPSLA 2026 · 1 citation
