Lune

NeurIPS2024顶会

Nimbus: Secure and Efficient Two-Party Inference for Transformers

Zhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu, Haoqi Wu, Xiao Wang, Yu Yu, Derun Zhao, Yancheng Zheng, Minyi Guo, Jingwen Leng

2024年份
34被引次数
13顶会引用

摘要

Transformer models have gained significant attention due to their power in machine learning tasks. Their extensive deployment has raised concerns about the potential leakage of sensitive information during inference. However, when being applied to Transformers, existing approaches based on secure two-party computation (2PC) bring about efficiency limitations in two folds: (1) resource-intensive matrix multiplications in linear layers, and (2) complex non-linear activation functions like GELU\mathsf{GELU} and Softmax\mathsf{Softmax}. This work presents a new two-party inference framework Nimbus\mathsf{Nimbus} for Transformer models. For the linear layer, we propose a new 2PC paradigm along with an encoding approach to securely compute matrix multiplications based on an outer-product insight, which achieves 2.9×∼12.5×2.9\times \sim 12.5\times performance improvements compared to the state-of-the-art (SOTA) protocol. For the non-linear layer, through a new observation of utilizing the input distribution, we propose an approach of low-degree polynomial approximation for GELU\mathsf{GELU} and Softmax\mathsf{Softmax}, which improves the performance of the SOTA polynomial approximation by 2.9×∼4.0×2.9\times \sim 4.0\times, where the average accuracy loss of our approach is 0.08% compared to the non-2PC inference without privacy. Compared with the SOTA two-party inference, Nimbus\mathsf{Nimbus} improves the end-to-end performance of inference by 2.7×∼4.7×2.7\times \sim 4.7\times across different network settings.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper13

问问它们各自怎么用它

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖