A New PPML Paradigm for Quantized Models
Tianpei Lu, Bingsheng Zhang, Xiaoyuan Zhang, Kui Ren
Abstract
—Model quantization has become a common practice in machine learning (ML) to improve efficiency and reduce computational/communicational overhead. However, adopting quantization in privacy-preserving machine learning (PPML) remains challenging due to the complex internal structure of quantized operators, which leads to inefficient protocols under the existing PPML frameworks. In this work, we propose a new PPML paradigm that is tailor-made for and can benefit from quantized models. Our main observation is that lookup tables can ignore the complex internal constructs of any functions which can be used to simplify the quantized operator evaluation. We view the model inference process as a sequence of quantized operators, and each operator is implemented by a lookup table. We then develop an efficient private lookup table evaluation protocol, and its online communication cost is only log n , where n is the size of the lookup table. On a single CPU core, our protocol can evaluate 2 26 tables with 8-bit input and 8-bit output per second. The resulting PPML framework for quantized models offers extremely fast online performance. The experimental results demonstrate that our quantization strategy achieves substantial speedups over SOTA PPML solutions, improving the online performance by 40 ∼ 60 × w.r.t. convolutional neural network (CNN) models, such as AlexNet, VGG16, and ResNet18, and by 10 ∼ 25 × w.r.t. large language models (LLMs), such as GPT-2, GPT-Neo, and Llama2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5239f259-b400-4e6e-9745-e4bddab52717Cited by top-tier papers3
- MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM InferenceWenxuan Zeng, Ye Dong, Jinjin Zhou, Jin Tan et al.NeurIPS 2025 · 4 citations
- Sort, Sweep, Mirror: Batch Private Interval Lookup with Logarithmic CostAndes Y. L. Kei, Lucien K. L. Ng, Jack P. K. Ma, Sherman S. M. ChowS&P 2026 · 2 citations
- Sok: Private Transformer-based Model InferenceYuntian Chen, Tianpei Lu, Zhanyong Tang, Bingsheng Zhang et al.USENIX Security 2026
Builds on33
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- ABY3: A Mixed Protocol Framework for Machine LearningPayman Mohassel, Peter RindalCCS 2018 · 898 citations
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta et al.NeurIPS 2021 · 573 citations
- MASCOT: Faster Malicious Arithmetic Secure Computation with Oblivious TransferMarcel Keller, Emmanuela Orsini, Peter SchollCCS 2016 · 487 citations
- XONN: XNOR-based Oblivious Deep Neural Network InferenceM. Sadegh Riazi, Mohammad Samragh, Hao Chen, Kim Laine et al.USENIX Security 2019 · 314 citations
Related papers
- FastQuery: Communication-efficient Embedding Table Query for Private LLMs inferenceChenqi Lin, Tianshi Xu, Zebin Yang, Runsheng Wang et al.DAC 2024
- Preprocessed Private Function Evaluation: Achieving Sublinear Online Complexity for Lookup TablesTanping Zhou, Xiaoyi Wang, Yi Qu, Wenchao Liu et al.CCS 2026
- MD-ML: Super Fast Privacy-Preserving Machine Learning for Malicious Security with a Dishonest MajorityBoshi Yuan, Shixuan Yang, Yongxiang Zhang, Ning Ding et al.USENIX Security 2024 · 22 citations
- Privacy-Preserving Embedding via Look-up Table Evaluation with Fully Homomorphic EncryptionJaeyun Kim, Saerom Park, Joohee Lee, Jung Hee CheonICML 2024 · 5 citations
- FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up TablesGunho Park, Hyeokjun Kwon, Jiwoo Kim, Jeongin Bae et al.HPCA 2025 · 10 citations
