BOLT: Privacy-Preserving, Accurate and Efficient Inference for Transformers
Qi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng, Thomas Schneider
Abstract
The advent of transformers has brought about significant advancements in traditional machine learning tasks. However, their pervasive deployment has raised concerns about the potential leakage of sensitive information during inference. Existing approaches using secure multiparty computation (MPC) face limitations when applied to transformers due to the extensive model size and resource-intensive matrix-matrix multiplications. In this paper, we present BOLT, a privacy-preserving inference framework for transformer models that supports efficient matrix multiplications and nonlinear computations. Combined with our novel machine learning optimizations, BOLT reduces the communication cost by 10.91×. Our evaluation on diverse datasets demonstrates that BOLT maintains comparable accuracy to floating-point models and achieves 4.8-9.5× faster inference across various network settings compared to the state-of-the-art system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3688281-0d1e-4e43-810f-3426b4f1fa45Cited by top-tier papers51
- Nimbus: Secure and Efficient Two-Party Inference for TransformersZhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu et al.NeurIPS 2024 · 34 citations
- MOAI: Module-Optimizing Architecture for Non-Interactive Secure Transformer InferenceLinru Zhang, Xiangning Wang, Sim Jun Jie, Zhicong Huang et al.ICLR 2026 · 24 citations
- PrivCirNet: Efficient Private Inference via Block Circulant TransformationTianshi Xu, Lemeng Wu, Runsheng Wang, Meng LiNeurIPS 2024 · 21 citations
- Compass: Encrypted Semantic Search with High AccuracyJinhao Zhu, Liana Patel, Matei Zaharia, Raluca Ada PopaOSDI 2025 · 14 citations
- Powerformer: Efficient and High-Accuracy Privacy-Preserving Language Model with Homomorphic EncryptionDongjin Park, Eunsang Lee, Joon-Woo LeeACL 2025 · 13 citations
Builds on27
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta et al.NeurIPS 2021 · 573 citations
- I-BERT: Integer-only BERT QuantizationSehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney et al.ICML 2021 · 439 citations
Related papers
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCTianshi Xu, Wen-jie Lu, Jiangrui Yu, Yi Chen et al.USENIX Security 2025
- BumbleBee: Secure Two-party Inference Framework for Large TransformersWen-jie Lu, Zhicong Huang, Zhen Gu, Jingyu Li et al.NDSS 2025
- Ditto: Quantization-aware Secure Inference of Transformers upon MPCHaoqi Wu, Wenjing Fang, Yancheng Zheng, Junming Ma et al.ICML 2024 · 17 citations
- Mosformer: Maliciously Secure Three-Party Inference Framework for Large TransformersKe Cheng, Yuheng Xia, Anxiao Song, Jiaxuan Fu et al.CCS 2025
- THOR: Secure Transformer Inference with Homomorphic EncryptionJungho Moon, Dongwoo Yoo, Xiaoqian Jiang, Miran KimCCS 2025 · 1 citation
