USENIX Security2026Top-tier venue
TensorZKP: Repurposing GPU Tensor Cores for High-Performance Zero-Knowledge Proofs
Tao Lu, Jipeng Zhang, Yanpei Guo, Xuanming Liu, Wenjie Qu, Zonghui Wang, Wenzhi Chen, Jiaheng Zhang
Abstract
GPU Tensor Cores, specialized hardware units designed to accelerate matrix multiplication, have served as the primary engine behind the AI revolution. Given the exponential performance gains they have delivered, aligning cryptographic implementations with this hardware evolution is critical. This is particularly acute for zero-knowledge proofs (ZKPs), a cryptographic primitive that currently grapples with high proof generation costs. Existing GPU implementations for ZKPs rely exclusively on general-purpose SIMT cores, leaving the massive computational power of Tensor Cores untapped. In this paper, we introduce TensorZKP, the first GPU framework to harness Tensor Cores for ZKP acceleration. Since Tensor Cores are designed for low-precision matrix multiplication, mapping ZKP's arithmetic to this hardware is non-trivial. To bridge this gap, we develop Tensor-Core-compatible finite field arithmetic and reformulate ZKP modules, specifically sum-check protocols and Spielman code, into matrix multiplication tasks. Furthermore, we design an asynchronous warp-specialized framework that pipelines memory access, Tensor Core matrix operations, and SIMT-based modular reductions. We instantiate these optimizations with HyperPlonk as the Polynomial Interactive Oracle Proof (PIOP) and Brakedown as the Polynomial Commitment Scheme (PCS) to enable end-to-end proof generation. The evaluation results show that TensorZKP exhibits remarkable efficiency. At a 2^25 scale, the underlying building blocks complete in 0.85 ms for inner product, 0.91 ms for scalar-vector multiplication, 4.04 ms for degree-2 sum-check, and 11.58 ms for the encoder. For a circuit with 2^25 multiplication gates, TensorZKP achieves a proof generation time of only 215.28 milliseconds, representing a 955× speedup over the CPU baseline and a 36.2× improvement over state-of-the-art SIMT-based GPU implementations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49b3c94c-8056-4c1b-a6dc-80d5f2622c3eBuilds on27
- Doubly-Efficient zkSNARKs Without Trusted SetupRiad S. Wahby, Ioanna Tzialla, Abhi Shelat, Justin Thaler et al.S&P 2018 · 356 citations
- Marlin: Preprocessing zkSNARKs with Universal and Updatable SRSAlessandro Chiesa, Yuncong Hu, Mary Maller, Pratyush Mishra et al.EUROCRYPT 2020 · 356 citations
- Transparent SNARKs from DARK CompilersBenedikt Bünz, Ben Fisch, Alan SzepieniecEUROCRYPT 2020 · 240 citations
- Transparent Polynomial Delegation and Its Applications to Zero Knowledge ProofJiaheng Zhang, Tiancheng Xie, Yupeng Zhang, Dawn SongS&P 2020 · 192 citations
- DIZK: A Distributed Zero Knowledge Proof SystemHoward Wu, Wenting Zheng, Alessandro Chiesa, Raluca Ada Popa et al.USENIX Security 2018 · 152 citations
Related papers
- BatchZK: A Fully Pipelined GPU-Accelerated System for Batch Generation of Zero-Knowledge ProofsTao Lu, Yuxun Chen, Zonghui Wang, Xiaohang Wang et al.ASPLOS 2025 · 11 citations
- GenZA: A General and Efficient Accelerator for Diverse Zero-Knowledge Proof ProtocolsCheng Wang, Jiangbin Dong, Mingyu GaoISCA 2026
- Pipelonk: Accelerating End-to-End Zero-Knowledge Proof Generation on GPUs for PLONK-Based ProtocolsZhiyuan Zhang, Yanxin Cai, Wenhao Yin, Xueyu Wu et al.PPoPP 2026 · 1 citation
- Need for zkSpeed: Accelerating HyperPlonk for Zero-Knowledge ProofsAlhad Daftardar, Jianqiao Mo, Joey Ah-kiow, Benedikt Bünz et al.ISCA 2025 · 12 citations
- UniZK: Accelerating Zero-Knowledge Proof with Unified Hardware and Flexible Kernel MappingCheng Wang, Mingyu GaoASPLOS 2025 · 12 citations
