Ditto: Quantization-aware Secure Inference of Transformers upon MPC
Haoqi Wu, Wenjing Fang, Yancheng Zheng, Junming Ma, Jin Tan, Lei Wang
Abstract
Due to the rising privacy concerns on sensitive client data and trained models like Transformers, secure multi-party computation (MPC) techniques are employed to enable secure inference despite attendant overhead. Existing works attempt to reduce the overhead using more MPC-friendly non-linear function approximations. However, the integration of quantization widely used in plaintext inference into the MPC domain remains unclear. To bridge this gap, we propose the framework named Ditto to enable more efficient quantization-aware secure Transformer inference. Concretely, we first incorporate an MPC-friendly quantization into Transformer inference and employ a quantization-aware distillation procedure to maintain the model utility. Then, we propose novel MPC primitives to support the type conversions that are essential in quantization and implement the quantization-aware MPC execution of secure quantized inference. This approach significantly decreases both computation and communication overhead, leading to improvements in overall efficiency. We conduct extensive experiments on Bert and GPT2 models to evaluate the performance of Ditto. The results demonstrate that Ditto is about faster than MPCFormer (ICLR 2023) and faster than the state-of-the-art work PUMA with negligible utility degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4079d1b6-7dfa-4e83-bd14-5d6d4f4c6b1aCited by top-tier papers6
- Nimbus: Secure and Efficient Two-Party Inference for TransformersZhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu et al.NeurIPS 2024 · 34 citations
- FicGCN: Unveiling the Homomorphic Encryption Efficiency from Irregular Graph Convolutional NetworksZhaoxuan Kan, Husheng Han, Shangyi Shi, Tenghui Hua et al.ICML 2025
- Mosformer: Maliciously Secure Three-Party Inference Framework for Large TransformersKe Cheng, Yuheng Xia, Anxiao Song, Jiaxuan Fu et al.CCS 2025
- Cape: Context-Aware Prompt Perturbation Mechanism with Differential PrivacyHaoqi Wu, Wei Dai, Li Wang, Qiang YanICML 2025
- ObCLIP: Oblivious CLoud-Device Hybrid Image Generation with Privacy PreservationHaoqi Wu, Wei Dai, Ming Xu, Li Wang et al.NeurIPS 2025
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
Related papers
- BOLT: Privacy-Preserving, Accurate and Efficient Inference for TransformersQi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng et al.S&P 2024 · 149 citations
- MPCViT: Searching for Accurate and Efficient MPC-Friendly Vision Transformer with Heterogeneous AttentionWenxuan Zeng, Meng Li, Wenjie Xiong, Tong Tong et al.ICCV 2023 · 38 citations
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCTianshi Xu, Wen-jie Lu, Jiangrui Yu, Yi Chen et al.USENIX Security 2025
- PCFormer: Accelerating Privacy-preserving Transformer Inference by Partition and CombinationBo Zeng, Zhi Pang, Yuyang Zhang, Kai Zhao et al.AAAI 2026
- Privacy-Friendly Adaptation of Vision Transformers for Communication and Latency-Efficient Private InferenceZhi Pang, Bo Feng, Meng Luo, Chenhao Liu et al.WWW 2026
