TensorZKP: Repurposing GPU Tensor Cores for High-Performance Zero-Knowledge Proofs
Tao Lu, Jipeng Zhang, Yanpei Guo, Xuanming Liu, Wenjie Qu, Zonghui Wang, Wenzhi Chen, Jiaheng Zhang
摘要
GPU Tensor Cores, specialized hardware units designed to accelerate matrix multiplication, have served as the primary engine behind the AI revolution. Given the exponential performance gains they have delivered, aligning cryptographic implementations with this hardware evolution is critical. This is particularly acute for zero-knowledge proofs (ZKPs), a cryptographic primitive that currently grapples with high proof generation costs. Existing GPU implementations for ZKPs rely exclusively on general-purpose SIMT cores, leaving the massive computational power of Tensor Cores untapped. In this paper, we introduce TensorZKP, the first GPU framework to harness Tensor Cores for ZKP acceleration. Since Tensor Cores are designed for low-precision matrix multiplication, mapping ZKP's arithmetic to this hardware is non-trivial. To bridge this gap, we develop Tensor-Core-compatible finite field arithmetic and reformulate ZKP modules, specifically sum-check protocols and Spielman code, into matrix multiplication tasks. Furthermore, we design an asynchronous warp-specialized framework that pipelines memory access, Tensor Core matrix operations, and SIMT-based modular reductions. We instantiate these optimizations with HyperPlonk as the Polynomial Interactive Oracle Proof (PIOP) and Brakedown as the Polynomial Commitment Scheme (PCS) to enable end-to-end proof generation. The evaluation results show that TensorZKP exhibits remarkable efficiency. At a 2^25 scale, the underlying building blocks complete in 0.85 ms for inner product, 0.91 ms for scalar-vector multiplication, 4.04 ms for degree-2 sum-check, and 11.58 ms for the encoder. For a circuit with 2^25 multiplication gates, TensorZKP achieves a proof generation time of only 215.28 milliseconds, representing a 955× speedup over the CPU baseline and a 36.2× improvement over state-of-the-art SIMT-based GPU implementations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Doubly-Efficient zkSNARKs Without Trusted SetupRiad S. Wahby, Ioanna Tzialla, Abhi Shelat, Justin Thaler 等S&P 2018 · 被引用 356 次
- Marlin: Preprocessing zkSNARKs with Universal and Updatable SRSAlessandro Chiesa, Yuncong Hu, Mary Maller, Pratyush Mishra 等EUROCRYPT 2020 · 被引用 356 次
- Transparent SNARKs from DARK CompilersBenedikt Bünz, Ben Fisch, Alan SzepieniecEUROCRYPT 2020 · 被引用 240 次
- Transparent Polynomial Delegation and Its Applications to Zero Knowledge ProofJiaheng Zhang, Tiancheng Xie, Yupeng Zhang, Dawn SongS&P 2020 · 被引用 192 次
- DIZK: A Distributed Zero Knowledge Proof SystemHoward Wu, Wenting Zheng, Alessandro Chiesa, Raluca Ada Popa 等USENIX Security 2018 · 被引用 152 次
相关 Paper
- BatchZK: A Fully Pipelined GPU-Accelerated System for Batch Generation of Zero-Knowledge ProofsTao Lu, Yuxun Chen, Zonghui Wang, Xiaohang Wang 等ASPLOS 2025 · 被引用 11 次
- GenZA: A General and Efficient Accelerator for Diverse Zero-Knowledge Proof ProtocolsCheng Wang, Jiangbin Dong, Mingyu GaoISCA 2026
- Pipelonk: Accelerating End-to-End Zero-Knowledge Proof Generation on GPUs for PLONK-Based ProtocolsZhiyuan Zhang, Yanxin Cai, Wenhao Yin, Xueyu Wu 等PPoPP 2026 · 被引用 1 次
- Need for zkSpeed: Accelerating HyperPlonk for Zero-Knowledge ProofsAlhad Daftardar, Jianqiao Mo, Joey Ah-kiow, Benedikt Bünz 等ISCA 2025 · 被引用 12 次
- UniZK: Accelerating Zero-Knowledge Proof with Unified Hardware and Flexible Kernel MappingCheng Wang, Mingyu GaoASPLOS 2025 · 被引用 12 次
