Lune

USENIX Security2026Top-tier venue

TensorZKP: Repurposing GPU Tensor Cores for High-Performance Zero-Knowledge Proofs

Tao Lu, Jipeng Zhang, Yanpei Guo, Xuanming Liu, Wenjie Qu, Zonghui Wang, Wenzhi Chen, Jiaheng Zhang

2026Year

Abstract

GPU Tensor Cores, specialized hardware units designed to accelerate matrix multiplication, have served as the primary engine behind the AI revolution. Given the exponential performance gains they have delivered, aligning cryptographic implementations with this hardware evolution is critical. This is particularly acute for zero-knowledge proofs (ZKPs), a cryptographic primitive that currently grapples with high proof generation costs. Existing GPU implementations for ZKPs rely exclusively on general-purpose SIMT cores, leaving the massive computational power of Tensor Cores untapped. In this paper, we introduce TensorZKP, the first GPU framework to harness Tensor Cores for ZKP acceleration. Since Tensor Cores are designed for low-precision matrix multiplication, mapping ZKP's arithmetic to this hardware is non-trivial. To bridge this gap, we develop Tensor-Core-compatible finite field arithmetic and reformulate ZKP modules, specifically sum-check protocols and Spielman code, into matrix multiplication tasks. Furthermore, we design an asynchronous warp-specialized framework that pipelines memory access, Tensor Core matrix operations, and SIMT-based modular reductions. We instantiate these optimizations with HyperPlonk as the Polynomial Interactive Oracle Proof (PIOP) and Brakedown as the Polynomial Commitment Scheme (PCS) to enable end-to-end proof generation. The evaluation results show that TensorZKP exhibits remarkable efficiency. At a 2^25 scale, the underlying building blocks complete in 0.85 ms for inner product, 0.91 ms for scalar-vector multiplication, 4.04 ms for degree-2 sum-check, and 11.58 ms for the encoder. For a circuit with 2^25 multiplication gates, TensorZKP achieves a proof generation time of only 215.28 milliseconds, representing a 955× speedup over the CPU baseline and a 36.2× improvement over state-of-the-art SIMT-based GPU implementations.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 49b3c94c-8056-4c1b-a6dc-80d5f2622c3e

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines