PCFormer: Accelerating Privacy-preserving Transformer Inference by Partition and Combination
Bo Zeng, Zhi Pang, Yuyang Zhang, Kai Zhao, Tian Wu, Geying Yang, Lina Wang, Run Wang
Abstract
In recent years, transformer-based models have achieved remarkable success in sensitive domains, including healthcare, finance and personalized services, but their deployment raises significant privacy concerns. Existing secure inference studies have introduced cryptographic techniques such as Homomorphic Encryption (HE) and Secure Multi-Party Computation (MPC). However, these approaches either target isolated model components or incur prohibitive computational and communication overheads, failing to support latency-sensitive or resource-limited environments. In our investigation, we identify substantial redundancy in the nonlinear operations and their alternation with linear layers in deep learning. Motivated by this observation, we propose PCFormer, a universal optimization methodology tailored for sequences of linear and nonlinear computations in the Transformer. PCFormer introduces structure-aware partition and combination techniques specially designed for Multi-Head Attention (MHA) and Feed-Forward Network (FFN). Specifically, we reveal the discrete sources of redundancy in the Softmax and GeLU functions during inference, implementing partitions at the token and channel levels, respectively. Subsequently, these reductions are then combined with the preceding and succeeding linear operations, thereby enhancing both computational and communication efficiency. Experimental results on GLUE benchmarks demonstrate that PCFormer achieves a 1.9× speedup in both computation and communication without compromising accuracy, compared to existing privacy-preserving Transformer frameworks. Furthermore, we demonstrate that PCFormer generalizes effectively to other deep learning architectures involving structured linear-nonlinear compositions under cryptographic constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94ee50ec-ebdb-41c2-a45f-0944feb4b77cBuilds on12
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing et al.NeurIPS 2022 · 209 citations
- BOLT: Privacy-Preserving, Accurate and Efficient Inference for TransformersQi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng et al.S&P 2024 · 149 citations
- NumHTML: Numeric-Oriented Hierarchical Transformer Model for Multi-Task Financial ForecastingLinyi Yang, Jiazheng Li, Ruihai Dong, Yue Zhang et al.AAAI 2022 · 54 citations
- Nimbus: Secure and Efficient Two-Party Inference for TransformersZhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu et al.NeurIPS 2024 · 34 citations
- PrivCirNet: Efficient Private Inference via Block Circulant TransformationTianshi Xu, Lemeng Wu, Runsheng Wang, Meng LiNeurIPS 2024 · 21 citations
Related papers
- Powerformer: Efficient and High-Accuracy Privacy-Preserving Language Model with Homomorphic EncryptionDongjin Park, Eunsang Lee, Joon-Woo LeeACL 2025 · 13 citations
- MPCViT: Searching for Accurate and Efficient MPC-Friendly Vision Transformer with Heterogeneous AttentionWenxuan Zeng, Meng Li, Wenjie Xiong, Tong Tong et al.ICCV 2023 · 38 citations
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCTianshi Xu, Wen-jie Lu, Jiangrui Yu, Yi Chen et al.USENIX Security 2025
- Mosformer: Maliciously Secure Three-Party Inference Framework for Large TransformersKe Cheng, Yuheng Xia, Anxiao Song, Jiaxuan Fu et al.CCS 2025
- THOR: Secure Transformer Inference with Homomorphic EncryptionJungho Moon, Dongwoo Yoo, Xiaoqian Jiang, Miran KimCCS 2025 · 1 citation
