Primer: Fast Private Transformer Inference on Encrypted Data
Mengxin Zheng, Qian Lou, Lei Jiang
Abstract
It is increasingly important to enable privacy-preserving inference for cloud services based on Transformers. Post-quantum cryptographic techniques, e.g., fully homomorphic encryption (FHE), and multi-party computation (MPC), are popular methods to support private Transformer inference. However, existing works still suffer from prohibitively computational and communicational overhead. In this work, we present, Primer, to enable a fast and accurate Transformer over encrypted data for natural language processing tasks. In particular, Primer is constructed by a hybrid cryptographic protocol optimized for attention-based Transformer models, as well as techniques including computation merge and tokens-first ciphertext packing. Comprehensive experiments on encrypted language modeling show that Primer achieves state-of-the-art accuracy and reduces the inference latency by 90.6% ∼ 97.5% over previous methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f97aecca-f1ef-4bb9-b5dd-88d227691114Cited by top-tier papers10
- PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and RestorationZiqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu et al.ACL 2025 · 31 citations
- Converting Transformers to Polynomial Form for Secure Inference Over Homomorphic EncryptionItamar Zimerman, Moran Baruch, Nir Drucker, Gilad Ezov et al.ICML 2024 · 26 citations
- CENTAUR: Bridging the Impossible Trinity of Privacy, Efficiency, and Performance in Privacy-Preserving Transformer InferenceJinglong Luo, Guanzhong Chen, Yehong Zhang, Shiyu Liu et al.ACL 2025 · 9 citations
- DictPFL: Efficient and Private Federated Learning on Encrypted GradientsJiaqi Xue, Mayank Kumar, Yuzhang Shang, Shangqian Gao et al.NeurIPS 2025 · 4 citations
- zkVC: Fast Zero-Knowledge Proof for Private and Verifiable ComputingYancheng Zhang, Mengxin Zheng, Xun Chen, Jingtong Hu et al.DAC 2025 · 3 citations
Builds on5
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- EVA: an encrypted vector arithmetic language and compiler for efficient homomorphic computationRoshan Dathathri, Blagovesta Kostova, Olli Saarikivi, Wei Dai et al.PLDI 2020 · 117 citations
- SAFENet: A Secure, Accurate and Fast Neural Network InferenceQian Lou, Yilin Shen, Hongxia Jin, Lei JiangICLR 2021 · 65 citations
- AutoPrivacy: Automated Layer-wise Parameter Selection for Secure Neural Network InferenceQian Lou, Song Bian, Lei JiangNeurIPS 2020 · 41 citations
- Delphi: A Cryptographic Inference Service for Neural NetworksPratyush Mishra, Ryan Lehmkuhl, Akshayaram Srinivasan, Wenting Zheng et al.USENIX Security 2020
Related papers
- MERGE: Fast Private Text GenerationZi Liang, Pinghui Wang, Ruofei Zhang, Nuo Xu et al.AAAI 2024 · 15 citations
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCTianshi Xu, Wen-jie Lu, Jiangrui Yu, Yi Chen et al.USENIX Security 2025
- CipherPrune: Efficient and Scalable Private Transformer InferenceYancheng Zhang, Jiaqi Xue, Mengxin Zheng, Mimi Xie et al.ICLR 2025
- Tricycle: Private Transformer Inference with Tricyclic EncodingsLawrence Lim, Vikas Kalagi, Julia Novick, Jiaming Liu et al.CCS 2026 · 10 citations
- Encryption-Friendly LLM ArchitectureDonghwan Rho, Taeseong Kim, Minje Park, Jung Woo Kim et al.ICLR 2025
