Lune

HPCA2026顶会

HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation

Yaoyun Zhou, Qian Wang

2026年份

摘要

SPHINCS+is a stateless hash-based signature scheme known for its strong post-quantum security, but it suffers from slow signature speed due to the heavy use of hash operations. The parallel architecture of GPUs offers a potential advantage for accelerating the computation of SPHINCS+signatures. However, existing GPU-based optimization efforts for SPHINCS+ either do not fully exploit the inherent parallelism of its Merkle Tree-based structure, or lack fine-grained, compilerlevel customization tailored to its diverse computational kernels. This paper proposes HERO-Sign, which adopts hierarchical tuning methodologies and efficient compiler-time GPU optimizations for SPHINCS+. HERO-Sign rethinks the parallelization potential arising from data independence in SPHINCS+, components, including FORS (Forest of Random Subsets), MSS (Merkle Signature Scheme, and WOTS+(Winternitz One-Time Signature Plus). First, it introduces a Tree Fusion strategy for FORS, whose structure contains a large number of branches. Our FORS Fusion strategy is supported by an automated Tree Tuning search algorithm, allowing it to adapt and optimize fusion schemes across various GPU platforms. To further enhance performance, HERO-Sign adopts an adaptive compilation strategy that accounts for the varying effectiveness of compiler optimizations across different SPHINCS+component kernels (FORS_Sign, TREE_Sign, WOTS+_Sign). This strategy automatically selects between PTX and native branches during the compilation phase to maximize efficiency. For multiple batches of message signatures, HERO-Sign focuses on optimizing kernel-level overlapping and employs a Task Graph-based construction strategy to minimize multi-stream idle time and reduce kernel launch overhead. Compared to state-of-the-art GPU implementations, under the SPHINCS+-128f, SPHINCS+-192f and SPHINCS+256f parameter sets, HERO-Sign demonstrates an enhanced throughput of1.28×−3.13×, 1.28×−2.92×\mathbf{1. 2 8} \times \mathbf{- 3. 1 3} \times \mathbf{, ~} \mathbf{1. 2 8} \times \mathbf{- 2. 9 2} \times, and1.24×−2.60×\mathbf{1. 2 4} \times \mathbf{- 2. 6 0} \timeson RTX 4090. Similar performance improvements have also been achieved on other architectures, including the A100, H100, and GTX 2080. HERO-Sign also achieves a two-order-of-magnitude reduction in kernel launch latency.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖