HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
Yaoyun Zhou, Qian Wang
摘要
SPHINCS+is a stateless hash-based signature scheme known for its strong post-quantum security, but it suffers from slow signature speed due to the heavy use of hash operations. The parallel architecture of GPUs offers a potential advantage for accelerating the computation of SPHINCS+signatures. However, existing GPU-based optimization efforts for SPHINCS+ either do not fully exploit the inherent parallelism of its Merkle Tree-based structure, or lack fine-grained, compilerlevel customization tailored to its diverse computational kernels. This paper proposes HERO-Sign, which adopts hierarchical tuning methodologies and efficient compiler-time GPU optimizations for SPHINCS+. HERO-Sign rethinks the parallelization potential arising from data independence in SPHINCS+, components, including FORS (Forest of Random Subsets), MSS (Merkle Signature Scheme, and WOTS+(Winternitz One-Time Signature Plus). First, it introduces a Tree Fusion strategy for FORS, whose structure contains a large number of branches. Our FORS Fusion strategy is supported by an automated Tree Tuning search algorithm, allowing it to adapt and optimize fusion schemes across various GPU platforms. To further enhance performance, HERO-Sign adopts an adaptive compilation strategy that accounts for the varying effectiveness of compiler optimizations across different SPHINCS+component kernels (FORS_Sign, TREE_Sign, WOTS+_Sign). This strategy automatically selects between PTX and native branches during the compilation phase to maximize efficiency. For multiple batches of message signatures, HERO-Sign focuses on optimizing kernel-level overlapping and employs a Task Graph-based construction strategy to minimize multi-stream idle time and reduce kernel launch overhead. Compared to state-of-the-art GPU implementations, under the SPHINCS+-128f, SPHINCS+-192f and SPHINCS+256f parameter sets, HERO-Sign demonstrates an enhanced throughput of, andon RTX 4090. Similar performance improvements have also been achieved on other architectures, including the A100, H100, and GTX 2080. HERO-Sign also achieves a two-order-of-magnitude reduction in kernel launch latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- SPHINCS+C: Compressing SPHINCS+ With (Almost) No CostAndreas Hülsing, Mikhail A. Kudinov, Eyal Ronen, Eylon YogevS&P 2023
- Revisiting the Constant-Sum Winternitz One-Time Signature with Applications to SPHINCS+ and XMSSKaiyi Zhang, Hongrui Cui, Yu YuCRYPTO 2023 · 被引用 7 次
- Shorter Hash-Based Signatures Using Forced PruningMehdi Abri, Jonathan KatzCRYPTO 2026
- DGSP: An Efficient Scalable Fully Dynamic Group Signature Scheme Using rmSPHINCS+Mojtaba Fadavi, Seyyed Arash Azimi, Sabyasachi Karati, Samuel JaquesEUROCRYPT 2026
- HERO: A Hierarchical Set Partitioning and Join Framework for Speeding up the Set Intersection Over GraphsBoyu Yang, Weiguo Zheng, Xiang Lian, Yuzheng Cai 等SIGMOD 2024 · 被引用 5 次
