Why the Proof Fails in Different Versions of Theorem Provers: An Empirical Study of Compatibility Issues in Isabelle
Xiaokun Luan, David Sanán, Zhe Hou, Qiyuan Xu, Chengwei Liu, Yufan Cai, Yang Liu, Meng Sun
摘要
Proof assistants are software tools for formal modeling and verification of software, hardware, design, and mathematical proofs. Due to the growing complexity and scale of formal proofs, compatibility issues frequently arise when using different versions of proof assistants. These issues result in broken proofs, disrupting the maintenance of formalized theories and hindering the broader dissemination of results within the community. Although existing works have proposed techniques to address specific types of compatibility issues, the overall characteristics of these issues remain largely unexplored. To address this gap, we conduct the first extensive empirical study to characterize compatibility issues, using Isabelle as a case study. We develop a regression testing framework to automatically collect compatibility issues from the Archive of Formal Proofs, the largest repository of formal proofs in Isabelle. By analyzing 12,079 collected issues, we identify their types and symptoms and further investigate their root causes. We also extract updated proofs that address these issues to understand the applied resolution strategies. Our study provides an in-depth understanding of compatibility issues in proof assistants, offering insights that support the development of effective techniques to mitigate these issues.
CCS Concepts: • Software and its engineering → Maintaining software.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Thor: Wielding Hammers to Integrate Language Models and Automated Theorem ProversAlbert Qiaochu Jiang, Wenda Li, Szymon Tworkowski, Konrad Czechowski 等NeurIPS 2022 · 被引用 154 次
- LEGO-Prover: Neural Theorem Proving with Growing LibrariesHaiming Wang, Huajian Xin, Chuanyang Zheng, Zhengying Liu 等ICLR 2024 · 被引用 125 次
- Baldur: Whole-Proof Generation and Repair with Large Language ModelsEmily First, Markus N. Rabe, Talia Ringer, Yuriy BrunFSE 2023 · 被引用 89 次
- Has My Release Disobeyed Semantic Versioning? Static Detection Based on Semantic DifferencingLyuye Zhang, Chengwei Liu, Zhengzi Xu, Sen Chen 等ASE 2022 · 被引用 30 次
- Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal ProofsAlbert Qiaochu Jiang, Sean Welleck, Jin Peng Zhou, Timothée Lacroix 等ICLR 2023 · 被引用 25 次
相关 Paper
- Compatibility Issues in Deep Learning Systems: Problems and OpportunitiesJun Wang, Guanping Xiao, Shuai Zhang, Huashan Lei 等FSE 2023 · 被引用 13 次
- QED in Context: An Observation Study of Proof Assistant UsersJessica Shi, Cassia Torczon, Harrison Goldstein, Benjamin C. Pierce 等OOPSLA 2025 · 被引用 2 次
- Automatically detecting API-induced compatibility issues in Android apps: a comparative analysis (replicability study)Pei Liu, Yanjie Zhao, Haipeng Cai, Mattia Fazzini 等ISSTA 2022 · 被引用 21 次
- Formalising CXL Cache CoherenceChengsong Tan, Alastair F. Donaldson, John WickersonASPLOS 2025 · 被引用 12 次
- Towards Automatically Repairing Compatibility Issues in Published Android AppsYanjie Zhao, Li Li, Kui Liu, John C. GrundyICSE 2022 · 被引用 25 次
