HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
Yaoyun Zhou, Qian Wang
Abstract
SPHINCS+is a stateless hash-based signature scheme known for its strong post-quantum security, but it suffers from slow signature speed due to the heavy use of hash operations. The parallel architecture of GPUs offers a potential advantage for accelerating the computation of SPHINCS+signatures. However, existing GPU-based optimization efforts for SPHINCS+ either do not fully exploit the inherent parallelism of its Merkle Tree-based structure, or lack fine-grained, compilerlevel customization tailored to its diverse computational kernels. This paper proposes HERO-Sign, which adopts hierarchical tuning methodologies and efficient compiler-time GPU optimizations for SPHINCS+. HERO-Sign rethinks the parallelization potential arising from data independence in SPHINCS+, components, including FORS (Forest of Random Subsets), MSS (Merkle Signature Scheme, and WOTS+(Winternitz One-Time Signature Plus). First, it introduces a Tree Fusion strategy for FORS, whose structure contains a large number of branches. Our FORS Fusion strategy is supported by an automated Tree Tuning search algorithm, allowing it to adapt and optimize fusion schemes across various GPU platforms. To further enhance performance, HERO-Sign adopts an adaptive compilation strategy that accounts for the varying effectiveness of compiler optimizations across different SPHINCS+component kernels (FORS_Sign, TREE_Sign, WOTS+_Sign). This strategy automatically selects between PTX and native branches during the compilation phase to maximize efficiency. For multiple batches of message signatures, HERO-Sign focuses on optimizing kernel-level overlapping and employs a Task Graph-based construction strategy to minimize multi-stream idle time and reduce kernel launch overhead. Compared to state-of-the-art GPU implementations, under the SPHINCS+-128f, SPHINCS+-192f and SPHINCS+256f parameter sets, HERO-Sign demonstrates an enhanced throughput of, andon RTX 4090. Similar performance improvements have also been achieved on other architectures, including the A100, H100, and GTX 2080. HERO-Sign also achieves a two-order-of-magnitude reduction in kernel launch latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b966e7a3-24cf-4dcb-8d77-38142159245dBuilds on2
- The SPHINCS+ Signature FrameworkDaniel J. Bernstein, Andreas Hülsing, Stefan Kölbl, Ruben Niederhagen et al.CCS 2019 · 385 citations
- BatchZK: A Fully Pipelined GPU-Accelerated System for Batch Generation of Zero-Knowledge ProofsTao Lu, Yuxun Chen, Zonghui Wang, Xiaohang Wang et al.ASPLOS 2025 · 11 citations
Related papers
- SPHINCS+C: Compressing SPHINCS+ With (Almost) No CostAndreas Hülsing, Mikhail A. Kudinov, Eyal Ronen, Eylon YogevS&P 2023
- Revisiting the Constant-Sum Winternitz One-Time Signature with Applications to SPHINCS+ and XMSSKaiyi Zhang, Hongrui Cui, Yu YuCRYPTO 2023 · 7 citations
- Shorter Hash-Based Signatures Using Forced PruningMehdi Abri, Jonathan KatzCRYPTO 2026
- DGSP: An Efficient Scalable Fully Dynamic Group Signature Scheme Using rmSPHINCS+Mojtaba Fadavi, Seyyed Arash Azimi, Sabyasachi Karati, Samuel JaquesEUROCRYPT 2026
- HERO: A Hierarchical Set Partitioning and Join Framework for Speeding up the Set Intersection Over GraphsBoyu Yang, Weiguo Zheng, Xiang Lian, Yuzheng Cai et al.SIGMOD 2024 · 5 citations
