SysMoBench: Evaluating AI on Formally Specifying Complex Real-World Systems
Qian Cheng, Ruize Tang, Emilie Ma, Finn Hackett, Peiyang He, Yiming Su, Ivan Beschastnikh, Yu Huang, Xiaoxing Ma, Tianyin Xu
摘要
Formal models are essential to specifying large, complex computer systems and verifying their correctness, but are notoriously expensive to write and maintain. Recent advances in generative AI show promise in generating certain forms of specifications. However, existing work mostly targets small code, not complete systems. It is unclear whether AI can deal with realistic system artifacts, as this requires abstracting their complex behavioral properties into formal models. We present SYSMOBENCH, a benchmark that evaluates AI's ability to formally model large, complex systems. We focus on concurrent and distributed systems, which are keystones of today's critical computing infrastructures, encompassing operating systems and cloud infrastructure. We use TLA + , the de facto specification language for concurrent and distributed systems, though the benchmark can be extended to other specification languages. We address the primary challenge of evaluating AI-generated models by automating metrics like syntactic and runtime correctness, conformance to system code, and invariant correctness. SYSMOBENCH currently includes eleven diverse system artifacts: the Raft implementation of Etcd and Redis, the leader election of ZooKeeper, the Spinlock, Mutex, and Ringbuffer in Asterinas OS, etc., with more being added. SYSMOBENCH enables us to understand the capabilities and limitations of today's LLMs and agents, putting tools in this area on a firm footing and opening up promising new research directions. INTRODUCTION Formal models are essential to specifying computer systems and reasoning about their correctness. They provide a mathematical foundation to document and verify the design of complex systems, such as distributed protocols and concurrent algorithms (Lamport, 2002; Tasiran et al., 2003; Newcombe et al., 2015; Hackett et al., 2023b) . Recently, formal models are used to describe system implementations-system code that runs on user devices and in production environments. Such models, which we refer to as system models, enable verification of system code via comprehensive testing and model checking (Bornholt et al., 2021; Tang et al., 2024; Ouyang et al., 2025; Tang et al., 2025) . For example, system models of Apache ZooKeeper (a distributed coordination system) were used to detect deep bugs that violate system safety and verify their fixes (Ouyang et al., 2025 ). However, system models are notoriously expensive to write and maintain. Different from protocols and algorithms, system code contains low-level details, is more complex, and constantly evolves. Hence, synthesis of system models is an open challenge (e.g., TLAi+ Challenge (2025)). Recent advances in generative AI, represented by large language models (LLMs) and agentic techniques, show promise in generating function-level specifications, in the form of pre-and postconditions (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Exploring and Unleashing the Power of Large Language Models in Automated Code TranslationZhen Yang, Fang Liu, Zhongxing Yu, Jacky Wai Keung 等FSE 2024 · 被引用 72 次
- Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3James Bornholt, Rajeev Joshi, Vytautas Astrauskas, Brendan Cully 等SOSP 2021 · 被引用 63 次
- Enchanting Program Specification Synthesis by Large Language Models Using Static Analysis and Program VerificationCheng Wen, Jialun Cao, Jie Su, Zhiwu Xu 等CAV 2024 · 被引用 60 次
相关 Paper
- OSVBench: Benchmarking LLMs on Specification Generation Tasks for Operating System VerificationShangyu Li, Juyong Jiang, Tiancheng Zhao, Jiasi ShenAAAI 2026 · 被引用 10 次
- SysBench: Can LLMs Follow System Message?Yanzhao Qin, Tao Zhang, Tao Zhang, Yanjun Shen 等ICLR 2025
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable CodeLingfei Zeng, Fengdi Che, Xuhan Huang, Fei Ye 等ICLR 2026 · 被引用 8 次
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development PracticesJia Li, Hongyi Deng, Yiran Zhang, Kechi Zhang 等FSE 2026
