SHAFT: Secure, Handy, Accurate and Fast Transformer Inference
Andes Y. L. Kei, Sherman S. M. Chow
Abstract
—Adoption of transformer-based machine learning models is growing, raising concerns about sensitive data exposure. Nonetheless, current secure inference solutions incur substantial overhead due to their extensive reliance on non-linear protocols, such as softmax and Gaussian error linear unit (GELU). Driven by numerical stability needs, softmax approximations ( e.g. , NeurIPS 2021) typically extract the maximum element of an input vector, incurring logarithmic rounds (in the input length). Existing GELU protocols ( e.g. , S&P 2024) use piecewise approximations with high-degree polynomials that rely heavily on secure multiplications and comparisons, which are expensive. Such complexities also hinder model owners unfamiliar with cryptography from deploying their custom models easily. SHAFT, our proposed system, provides a secure, handy, accurate, and fast transformer inference framework for deployment. Highlights of our contributions include 1) the first constant-round softmax protocol for transformers, uniquely combining the benefits of input clipping and characteristics of ordinary differential equations, and 2) a highly accurate GELU protocol on a novel characterization designed for Fourier series approximation. Extending to broader contexts, our new protocols also apply to general neural networks that use softmax as the final layer and to transformer architectures with different activation functions. Remarkably, SHAFT outperforms state-of-the-art SIGMA (PETS 2024), which uses secret sharing, and BumbleBee (NDSS 2025), which additionally uses RLWE-based homomorphic encryption. More specifically, SHAFT minimizes communication by 25 - 41% and matches SIGMA’s running time while surpassing BumbleBee in running time by 4 . 6 - 5 . 3 × on LANs and 2 . 9 - 4 . 4 × on WANs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d6e399d-3c95-48fc-9b34-5bb941a631eeCited by top-tier papers15
- CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-TuningSaisai Xia, Wenhao Wang, Zihao Wang, Yuhui Zhang et al.NDSS 2026 · 4 citations
- Sort, Sweep, Mirror: Batch Private Interval Lookup with Logarithmic CostAndes Y. L. Kei, Lucien K. L. Ng, Jack P. K. Ma, Sherman S. M. ChowS&P 2026 · 2 citations
- Shared Spotlight Meridian: Distributed Sparse Pseudorandom Functions for Scalable Federated LearningYoulong Ding, Peihua Mai, Jingqi Zhang, Sherman S. M. Chow et al.S&P 2026 · 1 citation
- Efficient and High-Accuracy Secure Two-Party Protocols for a Class of Functions with Real-number InputsHao Guo, Zhaoqian Liu, Liqiang Peng, Shuaishuai Li et al.USENIX Security 2026 · 1 citation
- On the (In-)Security of the Shuffling Defense in the Transformer Secure InferenceZhengyi Li, Yakai Wang, Jingwen Leng, Kang Yang et al.ACL 2026
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta et al.NeurIPS 2021 · 573 citations
- MASCOT: Faster Malicious Arithmetic Secure Computation with Oblivious TransferMarcel Keller, Emmanuela Orsini, Peter SchollCCS 2016 · 487 citations
- CrypTFlow: Secure TensorFlow InferenceNishant Kumar, Mayank Rathee, Nishanth Chandran, Divya Gupta et al.S&P 2020 · 276 citations
Related papers
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing et al.NeurIPS 2022 · 209 citations
- BumbleBee: Secure Two-party Inference Framework for Large TransformersWen-jie Lu, Zhicong Huang, Zhen Gu, Jingyu Li et al.NDSS 2025
- Nimbus: Secure and Efficient Two-Party Inference for TransformersZhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu et al.NeurIPS 2024 · 34 citations
- Mosformer: Maliciously Secure Three-Party Inference Framework for Large TransformersKe Cheng, Yuheng Xia, Anxiao Song, Jiaxuan Fu et al.CCS 2025
- THOR: Secure Transformer Inference with Homomorphic EncryptionJungho Moon, Dongwoo Yoo, Xiaoqian Jiang, Miran KimCCS 2025 · 1 citation
