Fast and Accurate Homomorphic Softmax Evaluation
Wonhee Cho, Guillaume Hanrot, Taeseong Kim, Minje Park, Damien Stehlé
Abstract
Homomorphic encryption is one of the main solutions for building secure and privacy-preserving solutions for Machine Learning as a Service, a major challenge in a society where AI becomes more and more pervasive. This motivates the development of homomorphic algorithms for the main building blocks of AI, typically for the components of the various types of neural networks architectures. Among those components, we focus on the Softmax function, defined by Softmax(x) = exp(𝑥 𝑖 )/ 𝑛 𝑗=1 exp(𝑥 𝑗 ) 1≤𝑖 ≤𝑛 * Corresponding author. in 486s (single-thread CPU), corresponding to an amortized 0.06s per Softmax call. All Softmax calls of the 32-layers LLaMa large language model (7B version) with context length 128 on an RTX-6000 GPU take around 1.5 minutes, and the final Softmax call in dimension 32768 for token generation takes less than 3 seconds. This suggests that near-practicality may be accessible with dedicated hardware.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bdfc225-50bf-4a9a-a46e-d0f55e927c5dCited by top-tier papers3
- Hyperion: Private Token Sampling with Homomorphic EncryptionLawrence Lim, Jiaming Liu, Vikas Kalagi, Divyakant Agrawal et al.ACL 2026 · 1 citation
- EncryptedLLM: Privacy-Preserving Large Language Model Inference via GPU-Accelerated Fully Homomorphic EncryptionLeo de Castro, Daniel Escudero, Adya Agrawal, Antigoni Polychroniadou et al.ICML 2025
- ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE EvaluationJiangrui Yu, Baosheng Zhang, Liang Kong, Lin Ding et al.CCS 2026
Builds on6
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- Secure Outsourced Matrix Computation and Application to Neural NetworksXiaoqian Jiang, Miran Kim, Kristin E. Lauter, Yongsoo SongCCS 2018 · 359 citations
- Efficient Bootstrapping for Approximate Homomorphic Encryption with Non-sparse KeysJean-Philippe Bossuat, Christian Mouchet, Juan Ramón Troncoso-Pastoriza, Jean-Pierre HubauxEUROCRYPT 2021 · 179 citations
- High-Precision Bootstrapping for Approximate Homomorphic Encryption by Error Variance MinimizationYongwoo Lee, Joon-Woo Lee, Young-Sik Kim, Yongjune Kim et al.EUROCRYPT 2022 · 67 citations
- HETAL: Efficient Privacy-preserving Transfer Learning with Homomorphic EncryptionSeewoo Lee, Garam Lee, Jung Woo Kim, Junbum Shin et al.ICML 2023 · 52 citations
Related papers
- FALCON: A Fourier Transform Based Approach for Fast and Secure Convolutional Neural Network PredictionsShaohua Li, Kaiping Xue, Bin Zhu, Chenkai Ding et al.CVPR 2020
- Cerium: A Multi-GPU Framework for Terabyte-Scale Encrypted InferenceSiddharth Jayashankar, Joshua Kim, Michael B. Sullivan, Wenting Zheng et al.SOSP 2026
- Encryption-Friendly LLM ArchitectureDonghwan Rho, Taeseong Kim, Minje Park, Jung Woo Kim et al.ICLR 2025
- Falcon: Fast Spectral Inference on Encrypted DataQian Lou, Wen-jie Lu, Cheng Hong, Lei JiangNeurIPS 2020 · 50 citations
- SLOTHE : Lazy Approximation of Non-Arithmetic Neural Network Functions over Encrypted DataKevin Nam, Youyeon Joo, Seungjin Ha, Yunheung PaekUSENIX Security 2025
