MERGE: Fast Private Text Generation
Zi Liang, Pinghui Wang, Ruofei Zhang, Nuo Xu, Shuo Zhang, Lifeng Xing, Haitao Bai, Ziyang Zhou
Abstract
The drastic increase in language models' parameters has led to a new trend of deploying models in cloud servers, raising growing concerns about private inference for Transformer-based models. Existing two-party privacy-preserving techniques, however, only take into account natural language understanding (NLU) scenarios. Private inference in natural language generation (NLG), crucial for applications like translation and code completion, remains underexplored. In addition, previous privacy-preserving techniques suffer from convergence issues during model training and exhibit poor inference speed when used with NLG models due to the neglect of time-consuming operations in auto-regressive generations. To address these issues, we propose a fast private text generation framework for Transformer-based language models, namely MERGE. MERGE reuses the output hidden state as the word embedding to bypass the embedding computation and reorganize the linear operations in the Transformer module to accelerate the forward procedure. Extensive experiments show that MERGE achieves a 26.5x speedup to the vanilla encrypted model under the sequence length 512, and reduces 80% communication cost, with an up to 10x speedup to state-of-the-art approximated models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98531ccd-e6e7-40eb-9b51-e4d0d1eb0440Cited by top-tier papers9
- Converting Transformers to Polynomial Form for Secure Inference Over Homomorphic EncryptionItamar Zimerman, Moran Baruch, Nir Drucker, Gilad Ezov et al.ICML 2024 · 26 citations
- CENTAUR: Bridging the Impossible Trinity of Privacy, Efficiency, and Performance in Privacy-Preserving Transformer InferenceJinglong Luo, Guanzhong Chen, Yehong Zhang, Shiyu Liu et al.ACL 2025 · 9 citations
- Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic DataZi Liang, Qingqing Ye, Xuan Liu, Yanyun Wang et al.NeurIPS 2025 · 5 citations
- Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingYanyun Wang, Li LiuICCV 2025 · 1 citation
- An Efficient Private GPT Never Autoregressively DecodesZhengyi Li, Yue Guan, Kang Yang, Yu Feng et al.ICML 2025
Builds on6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta et al.NeurIPS 2021 · 573 citations
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing et al.NeurIPS 2022 · 209 citations
Related papers
- Primer: Fast Private Transformer Inference on Encrypted DataMengxin Zheng, Qian Lou, Lei JiangDAC 2023 · 26 citations
- BumbleBee: Secure Two-party Inference Framework for Large TransformersWen-jie Lu, Zhicong Huang, Zhen Gu, Jingyu Li et al.NDSS 2025
- BOLT: Privacy-Preserving, Accurate and Efficient Inference for TransformersQi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng et al.S&P 2024 · 149 citations
- CipherPrune: Efficient and Scalable Private Transformer InferenceYancheng Zhang, Jiaqi Xue, Mengxin Zheng, Mimi Xie et al.ICLR 2025
- SecMoE: Communication-Efficient Secure MoE Inference via Select-Then-ComputeBowen Shen, Yuyue Chen, Peng Yang, Bin Zhang et al.AAAI 2026
