Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity
Guang Yan, Yuhui Zhang, Zimu Guo, Lutan Zhao, Xiaojun Chen, Chen Wang, Wenhao Wang, Dan Meng, Rui Hou
Abstract
With the growing use of large language models (LLMs) hosted on cloud platforms to offer inference services, privacy concerns about the potential leakage of sensitive information are escalating. Secure Multi-Party Computation (MPC) is a promising solution to protect the privacy in LLM inference. However, MPC requires frequent inter-server communication, causing high performance overhead. Inspired by the prevalent activation sparsity of LLMs, where most neuron are not activated after non-linear activation functions, we propose an efficient private inference system, Comet. This system employs an accurate and fast predictor to predict the sparsity distribution of activation function output. Additionally, we introduce a new private inference protocol. It efficiently and securely avoids computations involving zero values by exploiting the spatial locality of the predicted sparsity distribution. While this computation-avoidance approach impacts the spatiotemporal continuity of KV cache entries, we address this challenge with a low-communication overhead cache refilling strategy that merges miss requests and incorporates a prefetching mechanism. Finally, we evaluate Comet on four common LLMs and compare it with six state-of-the-art private inference systems. Comet achieves a speedup and a communication reduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9af6938-5f59-4278-adda-6b6b307cd06fCited by top-tier papers5
- SecP-Tuning: Efficient Privacy-Preserving Prompt Tuning for Large Language Models via MPCJinglong Luo, Zhuo Zhang, Yehong Zhang, Shiyu Liu et al.ICLR 2026 · 5 citations
- MEPS: Privacy-preserving Edge-cloud Video Foundation Model Inference with Privacy ProtectabilitySiping Shi, Rui Lu, Dan Wang, Bihai ZhangUSENIX Security 2026
- ReLUPruner: Rethinking ReLU Importance with Taylor Expansion for Efficient Private InferenceZhenpeng Li, Jinshuo Liu, Xinyan Wang, Lina Wang et al.AAAI 2026
- PCFormer: Accelerating Privacy-preserving Transformer Inference by Partition and CombinationBo Zeng, Zhi Pang, Yuyang Zhang, Kai Zhao et al.AAAI 2026
- Sok: Private Transformer-based Model InferenceYuntian Chen, Tianpei Lu, Zhanyong Tang, Bingsheng Zhang et al.USENIX Security 2026
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
Related papers
- MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM InferenceWenxuan Zeng, Ye Dong, Jinjin Zhou, Jin Tan et al.NeurIPS 2025 · 4 citations
- Privacy-Friendly Adaptation of Vision Transformers for Communication and Latency-Efficient Private InferenceZhi Pang, Bo Feng, Meng Luo, Chenhao Liu et al.WWW 2026
- BumbleBee: Secure Two-party Inference Framework for Large TransformersWen-jie Lu, Zhicong Huang, Zhen Gu, Jingyu Li et al.NDSS 2025
- Reconstruction Attack-Resistant Inference Paradigm for LLM Cloud ServicesZipeng Ye, Wenjian Luo, Qi Zhou, Yubo TangAAAI 2026
- COUNTDOWN: Contextually Sparse Activation Filtering Out Unnecessary Weights in Down ProjectionJaewon Cheon, Pilsung KangEMNLP 2025
