Neural Interactive Proofs
Lewis Hammond, Sam Adam-Day
Abstract
We consider the problem of how a trusted, but computationally bounded agent (a 'verifier') can learn to interact with one or more powerful but untrusted agents ('provers') in order to solve a given task without being misled. More specifically, we study the case in which agents are represented using neural networks and refer to solutions of this problem as neural interactive proofs. First we introduce a unifying framework based on proververifier games (Anil et al., 2021) , which generalises previously proposed interaction 'protocols'. We then describe several new protocols for generating neural interactive proofs, and provide a (theoretical) comparison of both new and existing approaches. In so doing, we aim to create a foundation for future work on neural interactive proofs and their application in building safer AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d04c8913-eb46-4e56-8275-c9bc3a55aa43Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- Implicit Learning Dynamics in Stackelberg Games: Equilibria Characterization, Convergence Analysis, and Empirical StudyTanner Fiez, Benjamin Chasnov, Lillian J. RatliffICML 2020 · 144 citations
- Proof-of-Learning: Definitions and PracticeHengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud et al.S&P 2021 · 132 citations
- Scalable AI Safety via Doubly-Efficient DebateJonah Brown-Cohen, Geoffrey Irving, Georgios PiliourasICML 2024 · 42 citations
Related papers
- Neural Concept Verifier: Scaling Prover-Verifier Games via Concept EncodingsBerkant Turan, Suhrab Asadulla, David Steinmann, Kristian Kersting et al.ICML 2026
- Analyzing Learning-Based Networked Systems with Formal VerificationArnaud Dethise, Marco Canini, Nina NarodytskaINFOCOM 2021 · 11 citations
- Learning to Prove Theorems by Learning to Generate TheoremsMingzhe Wang, Jia DengNeurIPS 2020 · 60 citations
- Towards Interpretable Deep Reinforcement Learning with Human-Friendly PrototypesEoin M. Kenny, Mycal Tucker, Julie ShahICLR 2023
- Automated Verification of Soundness of DNN CertifiersAvaljot Singh, Yasmin Sarita, Charith Mendis, Gagandeep SinghOOPSLA 2025 · 3 citations
