Lune

NeurIPS2025Top-tier venue

Generating and Checking DNN Verification Proofs

Hai Duong, ThanhVu Nguyen, Matthew Dwyer

2025Year
9Citations
1Top-tier citations

Abstract

Deep Neural Networks (DNN) have emerged as an effective approach to implementing challenging subproblems. They are increasingly being used as components in critical transportation, medical, and military systems. However, like human-written software, DNNs may have flaws that can lead to unsafe system performance. To confidently deploy DNNs in such systems, strong evidence is needed that they do not contain such flaws. This has led researchers to explore the adaptation and customization of software verification approaches to the problem of neural network verification (NNV). Many dozens of NNV tools have been developed in recent years and as a field these techniques have matured to the point where realistic networks can be analyzed to detect flaws and to prove conformance with specifications. NNV tools are highly-engineered and complex may harbor flaws that cause them to produce unsound results. We identify commonalities in algorithmic approaches taken by NNV tools to define a verifier independent proof format-activation pattern tree proofs (APTP)-and design an algorithm for checking those proofs that is proven correct and optimized to enable scalable checking. We demonstrate that existing verifiers can efficiently generate APTP proofs, and that an APTPchecker significantly outperforms prior work on a benchmark of 16 neural networks and 400 NNV problems, and that it is robust to variation in APTP proof structure arising from different NNV tools.

APTPchecker is available at: https://github.com/dynaroars/APTPchecker.

However, despite the progress in algorithmic advances, a fundamental question remains: "How can we trust the results produced by DNN verification tools?" While existing tools emit counterexamples when properties are violated (i.e., SAT results), there is no mechanism to independently validate results when properties are proven to hold (i.e., UNSAT claims). Recent competitions such as VNN-COMP [4] have revealed correctness issues in multiple tools, including cases where a verifier incorrectly declared a property to be proven even when a counterexample exists. These errors are difficult to detect and debug due to the complexity of verifier implementations, which often exceed tens of thousands of lines of code and employ intricate optimization techniques, e.g., top of the line DNN verification tools such as αβ-CROWN [7] and NeuralSAT [9] have 20k SLOC implementations with complex algorithms that may harbor bugs. Without a mechanism to independently validate 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

verification results, correctness of DNN verification tools cannot be assured and therefore posing a serious obstacle to deploying DNNs in safety-critical domains.

To address this, we propose proof-producing DNN verification: an approach in which verifiers emit a formal proof object that encodes the reasoning steps behind the verification result, and a separate, minimal proof checker certifies the proof's validity. This paradigm, long established in classical logic and SAT solving [10,11,12], brings transparency, auditability, and trust to the verification process.

More specifically, we analyze the broad class of Branch-and-Bound (BaB) DNN verification algorithms and reveal that they share two commonalities: (1) they refine the abstractions they use by performing case splitting to reason about the different phases of neuron activation, and (2) within cases they perform reasoning steps that can be formulated within the broad class of mixed integer linear programming (MILP) problems. Based on these insights, we show that BaB DNN verification naturally emit activation pattern tree proofs (APTP), which are a compact representation of the reasoning steps performed by the verifier ( §3.1). We also define a verifier independent APTP format that can be efficiently generated on-the-fly during DNN verification ( §3.2). Finally, we resent the APTPchecker algorithm along with a suite of optimizations and implement an independent APTPchecker prototype tool that has a small-footprint (800 SLOC) and validates APTP proofs using standard MILP solving. This paper makes the following contributions:

• We identify commonalities in BaB DNN verification algorithms and show how they can be minimally extended to generate proofs of unsatisfiability ( §3.1).

• We define a verifier-independent, compact, and SMTLib [13]-based human-readable proof format, APTP, that captures the reasoning steps of BaB verifiers ( §3.2).

• We implement the APTPchecker tool to independently and efficiently check APTP proofs ( §4).

• We evaluate our work on a benchmark of 400 verification problems involving 16 networks, including large models (up to 1.7M parameters) ( §5), and demonstrate that APTP and APTPchecker are robust to variation in proof structure arising from different DNN verification algorithms.

It is important to note that our goal is not to create a new DNN verifier, but to verify the correctness of res

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7fb993af-b2a2-42e1-81b5-347f586b1137

Cited by top-tier papers1

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines