MOAI: Module-Optimizing Architecture for Non-Interactive Secure Transformer Inference
Linru Zhang, Xiangning Wang, Sim Jun Jie, Zhicong Huang, Jiahao Zhong, Huaxiong Wang, Pu Duan, Kwok-Yan Lam
摘要
Privacy concerns have been raised in Large Language Models (LLM) inference when models are deployed in Cloud Service Providers (CSP). Homomorphic encryption (HE) offers a promising solution by enabling secure inference directly over encrypted inputs. However, the high computational overhead of HE remains a major bottleneck. To address this challenge, we propose MOAI, an efficient HE-based, non-interactive framework for secure transformer inference. MOAI gains significant efficiency improvement from: (1) a novel evaluation flow that combines column and diagonal packing with consistent strategies across all layers, eliminating expensive format conversions. (2) rotation-free algorithms for Softmax and LayerNorm that significantly reduce the number of costly HE rotations, removing 2448 HE rotations in BERT-base inference. (3) Column packing removes rotations in plaintext–ciphertext matrix multiplications and interleaved batching further reduces the rotations in ciphertext–ciphertext matrix multiplications. MOAI uses at least 1.7x fewer HE rotations compared to the state-of-the-art works across all matrix multiplications of BERT-base. As a result, We achieve a 52.8% reduction in evaluation time compared to the state-of-the-art HE-based non-interactive secure transformer inference, THOR (Moon et al., CCS'25). We then apply MOAI on the Powerformer's framework and achieve a 55.7% reduction in evaluation time compared to Powerformer (Park et al., ACL'25), which approximates Softmax and LayerNorm with simpler functions in transformer and proposes HE-based non-interactive transformer inference. We report an amortized time of 2.36 minutes per input on a single GPU environment. We show the extendibility by applying MOAI in LLaMA-3-8B. Our implementation is publicly available as open source.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Hyperion: Private Token Sampling with Homomorphic EncryptionLawrence Lim, Jiaming Liu, Vikas Kalagi, Divyakant Agrawal 等ACL 2026 · 被引用 1 次
- ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE EvaluationJiangrui Yu, Baosheng Zhang, Liang Kong, Lin Ding 等CCS 2026
- Sok: Private Transformer-based Model InferenceYuntian Chen, Tianpei Lu, Zhanyong Tang, Bingsheng Zhang 等USENIX Security 2026
它引用的顶会 Paper14
- Efficient Bootstrapping for Approximate Homomorphic Encryption with Non-sparse KeysJean-Philippe Bossuat, Christian Mouchet, Juan Ramón Troncoso-Pastoriza, Jean-Pierre HubauxEUROCRYPT 2021 · 被引用 179 次
- BOLT: Privacy-Preserving, Accurate and Efficient Inference for TransformersQi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng 等S&P 2024 · 被引用 149 次
- PEGASUS: Bridging Polynomial and Non-polynomial Evaluations in Homomorphic EncryptionWen-jie Lu, Zhicong Huang, Cheng Hong, Yiping Ma 等S&P 2021 · 被引用 139 次
- Orion: A Fully Homomorphic Encryption Framework for Deep LearningAustin Ebel, Karthik Garimella, Brandon ReagenASPLOS 2025 · 被引用 40 次
- Converting Transformers to Polynomial Form for Secure Inference Over Homomorphic EncryptionItamar Zimerman, Moran Baruch, Nir Drucker, Gilad Ezov 等ICML 2024 · 被引用 26 次
相关 Paper
- THOR: Secure Transformer Inference with Homomorphic EncryptionJungho Moon, Dongwoo Yoo, Xiaoqian Jiang, Miran KimCCS 2025 · 被引用 1 次
- Powerformer: Efficient and High-Accuracy Privacy-Preserving Language Model with Homomorphic EncryptionDongjin Park, Eunsang Lee, Joon-Woo LeeACL 2025 · 被引用 13 次
- Tricycle: Private Transformer Inference with Tricyclic EncodingsLawrence Lim, Vikas Kalagi, Julia Novick, Jiaming Liu 等CCS 2026 · 被引用 10 次
- Encryption-Friendly LLM ArchitectureDonghwan Rho, Taeseong Kim, Minje Park, Jung Woo Kim 等ICLR 2025
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCTianshi Xu, Wen-jie Lu, Jiangrui Yu, Yi Chen 等USENIX Security 2025
