Secure Outlier-Aware Large Language Model Inference
Lifan Zhao, Zhixuan Fang
Abstract
Secure multiparty computation allows the client to secretly inference their sensitive inputs without acquiring the proprietary machine learning model weights. As the decoder-only transformer-based large language model becomes the popular paradigm, the desire of applying MPC in large language models is increasing. However, such inference usually leads to great amount of latency, which is due to nonlinear operations in the Transformer architecture. Recent works either focus on improving cryptographic primitives or re-architecting and re-training to make LLM MPC-friendly. We, on the other hand, observe that properly addressing outlier phenomena, which are unique yet universal properties existing across different LLMs, can effectively reduce the input domain and thereby design faster protocols for non-linear operations. Hence, we propose Secure Outlier-Aware Large Language Model Inference framework (SOAL), which accelerates the RMSNorm operation by nearly 2 , SiLU by , and Softmax by more than 5. SOAL maintains the same performance of the original model without any fine-tuning requirement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa69d26b-6f1c-4a70-9c87-8624a0d6d1cdBuilds on17
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
Related papers
- BOLT: Privacy-Preserving, Accurate and Efficient Inference for TransformersQi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng et al.S&P 2024 · 149 citations
- SecP-Tuning: Efficient Privacy-Preserving Prompt Tuning for Large Language Models via MPCJinglong Luo, Zhuo Zhang, Yehong Zhang, Shiyu Liu et al.ICLR 2026 · 5 citations
- PCFormer: Accelerating Privacy-preserving Transformer Inference by Partition and CombinationBo Zeng, Zhi Pang, Yuyang Zhang, Kai Zhao et al.AAAI 2026
- MPCViT: Searching for Accurate and Efficient MPC-Friendly Vision Transformer with Heterogeneous AttentionWenxuan Zeng, Meng Li, Wenjie Xiong, Tong Tong et al.ICCV 2023 · 38 citations
- Mosformer: Maliciously Secure Three-Party Inference Framework for Large TransformersKe Cheng, Yuheng Xia, Anxiao Song, Jiaxuan Fu et al.CCS 2025
