Towards Closing the Performance Gap for Cryptographic Kernels Between CPUs and Specialized Hardware
Naifeng Zhang, Sophia Fu, Franz Franchetti
Abstract
Specialized hardware like application-specific integrated circuits (ASICs) remains the primary accelerator type for cryptographic kernels based on large integer arithmetic. Prior work has shown that commodity and server-class GPUs can achieve near-ASIC performance for these workloads. However, achieving comparable performance on CPUs remains an open challenge. This work investigates the following question: How can we narrow the performance gap between CPUs and specialized hardware for key cryptographic kernels like basic linear algebra subprograms (BLAS) operations and the number theoretic transform (NTT)?
To this end, we develop an optimized scalar implementation of these kernels for x86 CPUs at the per-core level. We utilize SIMD instructions-specifically AVX2 and AVX-512-to further improve performance, achieving an average speedup of 38 times and 62 times over state-of-the-art CPU baselines for NTTs and BLAS operations, respectively. To narrow the gap further, we propose a small AVX-512 extension, dubbed multi-word extension (MQX), which delivers substantial speedup with only three new instructions and minimal proposed hardware modifications. MQX cuts the slowdown relative to ASICs to as low as 35 times on a single CPU core. Finally, we perform a roofline analysis to evaluate the peak performance achievable with MQX when scaled across an entire multi-core CPU. Our results show that, with MQX, top-tier servergrade CPUs can approach the performance of state-of-the-art ASICs for cryptographic workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- F1: A Fast and Programmable Accelerator for Fully Homomorphic EncryptionNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Srinivas Devadas et al.MICRO 2021 · 294 citations
- CraterLake: a hardware accelerator for efficient unbounded computation on encrypted dataNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Nathan Manohar et al.ISCA 2022 · 205 citations
- TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPUShengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou et al.HPCA 2023 · 90 citations
- GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic EncryptionKaustubh Shivdikar, Yuhui Bao, Rashmi Agrawal, Michael Tian Shen et al.MICRO 2023 · 46 citations
Related papers
- A scalable SIMD RISC-V based processor with customized vector extensions for CRYSTALS-kyberHuimin Li, Nele Mentens, Stjepan PicekDAC 2022 · 13 citations
- BP-NTT: Fast and Compact in-SRAM Number Theoretic Transform with Bit-Parallel Modular MultiplicationJingyao Zhang, Mohsen Imani, Elaheh SadrediniDAC 2023 · 23 citations
- CryptoPIM: In-memory Acceleration for Lattice-based Cryptographic HardwareHamid Nejatollahi, Saransh Gupta, Mohsen Imani, Tajana Simunic Rosing et al.DAC 2020 · 63 citations
- Towards ML-KEM & ML-DSA on OpenTitanAmin Abdulrahman, Felix Oberhansl, Hoang Nguyen Hien Pham, Jade Philipoom et al.S&P 2025
- ENG25519: Faster TLS 1.3 handshake using optimized X25519 and Ed25519Jipeng Zhang, Junhao Huang, Lirui Zhao, Donglong Chen et al.USENIX Security 2024 · 11 citations
