How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs
Liyan Chen, Yael Kalai, Zoe Xi
Abstract
As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions. A recent line of work focuses on verification via debate, a model of interactive proofs where two competing powerful provers, or AI models, debate each other to convince a weak verifier, or a human, of the correctness of their claim. However, debate assumes that the two AI models possess equal abilities and that one of them is truthful, which may not be realistic. In this work, we show how to avoid debate : we initiate the study of single-prover interactive proofs for AI safety. Prior results in single-prover interactive proofs do not immediately carry over to the AI safety setting because they do not work when the computation has access to an oracle, such as to human judgment or an external database such as the web. We present doubly-efficient single-prover interactive proofs for oracle-aided computations (also known as relativizing proofs), in the settings where (1) the computation is robust, in the sense that the output does not change if at most a small fraction of the answers to oracle queries are incorrect, or (2) the oracle is a low-degree polynomial. These results suggest that interactive verification is possible even without debate, under structured or noise-tolerant oracle access.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 010ad4c4-c523-4263-bcc4-144564f11083Builds on4
- Scalable AI Safety via Doubly-Efficient DebateJonah Brown-Cohen, Geoffrey Irving, Georgios PiliourasICML 2024 · 42 citations
- Polynomial Commitments from Lattices: Post-quantum Security, Fast Verification and Transparent SetupValerio Cini, Giulio Malavolta, Ngoc Khanh Nguyen, Hoeteck WeeCRYPTO 2024 · 10 citations
- Efficiently Batching Unambiguous Interactive ProofsBonnie Berger, Rohan Goyal, Matthew M. Hong, Yael Tauman KalaiFOCS 2025
- Neural Interactive ProofsLewis Hammond, Sam Adam-DayICLR 2025
Related papers
- Undermining Mental Proof: How AI Can Make Cooperation Harder by Making Thinking EasierZachary Wojtowicz, Simon DeDeoAAAI 2025 · 7 citations
- Agentic Verification of Software SystemsHaoxin Tu, Huan Zhao, Yahui Song, Mehtab Zafar et al.FSE 2026 · 1 citation
- RedDebate: Safer Responses Through Multi-Agent Red Teaming DebatesAli Asad, Stephen Obadinma, Radin Shayanfar, Xiaodan ZhuICML 2026 · 7 citations
- Zero-Knowledge IOPs with Linear-Time Prover and Polylogarithmic-Time VerifierJonathan Bootle, Alessandro Chiesa, Siqi LiuEUROCRYPT 2022 · 27 citations
- DuoAI: Fast, Automated Inference of Inductive Invariants for Verifying Distributed ProtocolsJianan Yao, Runzhou Tao, Ronghui Gu, Jason NiehOSDI 2022 · 50 citations
