Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
Xu Xu, Xin Li, Xingwei Qu, Jie Fu, Binhang Yuan
Abstract
Despite rapid advances in code generation, current Large Language Models (LLMs) still lack an essential capability for reliable and verifiable code generation: compositional reasoning across multifunction programs. To explore this potential and important gap, we introduce DAFNYCOMP, a benchmark designed to systematically evaluate LLMs on the generation of compositional specifications in Dafny. Unlike prior benchmarks that primarily target single-function annotation, DAFNYCOMP focuses on programs composed of multiple interacting functions with necessary data dependencies, requiring LLMs to produce specifications that ensure correctness across component boundaries. Our benchmark comprises 300 automatically synthesized programs, each carefully constructed by combining 2-5 originally independent functions in a chain-based manner through LLM-driven synthesis. We evaluate LLMs from five leading research groups that represent the current frontier of reasoning-centric AI, including the GPT, CLAUDE, GEMINI, DEEPSEEK, and QWEN families. Our results reveal a striking dichotomy: while LLMs achieve both high syntax correctness (>99%) and moderate verification rates (>58%) in prior single-function benchmarks, they exhibit degraded syntax correctness (95.67%) and a catastrophic verification failure (3.69%) in DAFNYCOMP's compositional tasks-a 92% performance gap. Even the most powerful LLM achieves only 7% verification at Pass@8, with most LLMs below 2%. Further analysis reveals that LLMs systematically fail at cross-functional reasoning through three primary failure modes: specification fragility (39.2%), implementation-proof misalignment (21.7%), and reasoning instability (14.1%). These failures clearly reveal the absence of compositional reasoning capabilities in current LLMs. DAFNYCOMP thus establishes a diagnostic benchmark for tracking progress in verifiable code generation with LLMs, highlighting that the path from local to compositional verification remains largely uncharted.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6ab90e1-d458-499b-b67b-4fb97114bbf6Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong et al.ICLR 2024 · 92 citations
- Enchanting Program Specification Synthesis by Large Language Models Using Static Analysis and Program VerificationCheng Wen, Jialun Cao, Jie Su, Zhiwu Xu et al.CAV 2024 · 60 citations
- DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning GraphZhehao Zhang, Jiaao Chen, Diyi YangNeurIPS 2024 · 42 citations
- Towards AI-Assisted Synthesis of Verified Dafny MethodsMd Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James NobleFSE 2024 · 26 citations
Related papers
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable CodeLingfei Zeng, Fengdi Che, Xuhan Huang, Fei Ye et al.ICLR 2026 · 8 citations
- DRPBench: Evaluating LLMs in Concurrent Code Comprehension via Fine-grained Data Race PredictionYuqi Guo, Siwei Wei, Yan CaiICML 2026
- TASE: Token Awareness and Structured Evaluation for Multilingual Language ModelsChenzhuo Zhao, Xinda Wang, Yue Huang, Junting Lu et al.AAAI 2026 · 1 citation
- QuanBench: Benchmarking Quantum Code Generation with Large Language ModelsXiaoyu Guo, Minggu Wang, Jianjun ZhaoASE 2025 · 5 citations
