DeepProve: Verifiable End-to-End Large Language Model Inference
Nicolas Gailly, Ismael Hishon-Rezaizadeh, Tianyi Liu, Nicholas Mainardi, Dimitrios Papadopoulos, Charalampos Papamanthou, Christodoulos Pappas, Shravan Srinivasan, Zack Youell, Yupeng Zhang
摘要
Large Language Models (LLMs) are frontier deep learning systems that have achieved remarkable success across a wide range of AI services. However, their substantial computational and memory requirements make them difficult to deploy and run on local hardware. Due to these resource requirements, users often rely on untrusted cloud infrastructure providers to perform model inference. However, outsourcing introduces the challenge of verifying that the returned output is the genuine result of the specified model. In this work, we present DeepProve, the first system to enable efficient end-to-end verification of full LLM inference (i.e., for all generated tokens of a prompt) on untrusted cloud servers using zeroknowledge proofs (ZKPs). In contrast, prior work either provides only a proof-of-concept partial implementation for a single token (zkGPT, USENIX'25), or focuses exclusively on specific components of the inference pipeline, such as Softmax (zkLLM, CCS'24). DeepProve achieves end-to-end verification by certifying the correctness of the output sequence rather than encoding the expensive inference computation in-circuit, an approach that would require either circuit size quadratic in the sequence length or costly in-circuit modelling of RAM operations. The core building blocks of DeepProve are sum-check protocol and lookup arguments, which enable efficient proof of correctness of all operators needed for GPT-2 and Gemma 3, such as multi-head attention and layer normalization for GPT-2, and grouped-query attention, root mean square normalization, and rotary positional embeddings for Gemma 3. Our evaluation shows that DeepProve can prove inference of GPT-2 and Gemma 3 at approximately 174 and 86 tokens per minute, respectively, which is 20 -60× faster than the state of the art, without any significant loss in accuracy. Verification takes only 1 to 3.7 seconds. By distributing proof computation across multiple nodes, DeepProve can further improve the prover time while reducing the memory requirements for individual machines. With distributed proving, DeepProve can scale the throughput to 1855 tokens per minute. Our work represents the first full system for end-to-end LLM inference verification, thus paving the way for secure and trustworthy AI services. * Authors are listed alphabetically. † Work done while affiliated with Lagrange Labs. Preliminaries We include here some necessary background information for quantization and LLMs. Cryptographic definitions are in Appendix A. Quantization We use standard affine quantization techniques to convert the model weights and activations from floating-point to integer representations. The standard formula [JKC + 18] to convert a floating-point value 𝑥 into a quantized integer value 𝑥 is 𝑥 = 𝑠 • (𝑥 -𝑧), where 𝑠 is the scale factor and 𝑧 is the zero-point. Linear operations [JKC + 18] and requantization. Consider the matrix operation X 4 = X 1 X 2 + X 3 , where ×𝑛 , and X 3 , X 4 ∈ R 𝑚×𝑛 represented as floating-point values. We denote entries of each matrix as 𝑥 𝛼 [𝑖, 𝑗] for 𝛼 ∈ 1, 2, 3, 4. The quantization relations for each entry in the matrices are given by: 𝑥 𝛼 [𝑖, 𝑗] = 𝑠 𝛼 𝑥 𝛼 [𝑖, 𝑗] -𝑧 𝛼 . From the definition of matrix multiplication with bias addition: 𝑠 4 𝑥 4 [𝑖, 𝑘] -𝑧 4 = ℓ 𝑗=1 𝑠 1 𝑥 1 [𝑖, 𝑗] -𝑧 1 𝑠 2 𝑥 2 [ 𝑗, 𝑘] -𝑧 2 +𝑠 3 𝑥 3 [𝑖, 𝑘]-𝑧 3 , and 𝑥 4 [𝑖, 𝑘] = 𝑧 4 +𝑀 ℓ 𝑗=1 𝑥 1 [𝑖, 𝑗]-𝑧 1 𝑥 2 [ 𝑗, 𝑘]-𝑧 2 + 𝐵 𝑥 3 [𝑘] -𝑧 3 , where 𝑀 := 𝑠 1 𝑠 2 𝑠 4 and 𝐵 := 𝑠 3 𝑠 4 . In practice, ML frameworks, like PyTorch [For25], set 𝑠 3 = 𝑠 1 𝑠 2 . Since, 𝑀 is in (0, 1), we set 𝑀 = 2 -𝑛 𝜖 for some small 𝜖 > 0 and integer 𝑛.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu 等ICLR 2024 · 被引用 281 次
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui 等NeurIPS 2024 · 被引用 206 次
相关 Paper
- zkGPT: An Efficient Non-interactive Zero-knowledge Proof Framework for LLM InferenceWenjie Qu, Yijun Sun, Xuanming Liu, Tao Lu 等USENIX Security 2025
- zkLLM: Zero Knowledge Proofs for Large Language ModelsHaochen Sun, Jason Li, Hongyang ZhangCCS 2024 · 被引用 26 次
- TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable InferenceJack Min Ong, Matthew Di Ferrante, Aaron Pazdera, Ryan Garner 等ICML 2025
- Hollow-LLM Attack: Computationally Trivial Weights in Zero-Knowledge Verification of LLM InferenceChen Gong, Beijie Liu, Mengyuan LiS&P 2026 · 被引用 2 次
- zkAgent: Verifiable LLM Agent Execution via One-Shot Transcript ProofsLizheng Wang, Hancheng Lou, Chongrong Li, Yu Yu 等CCS 2026
