Selective Network Linearization for Efficient Private Inference
Minsu Cho, Ameya Joshi, Brandon Reagen, Siddharth Garg, Chinmay Hegde
摘要
Private inference (PI) enables inference directly on cryptographically secure data.While promising to address many privacy issues, it has seen limited use due to extreme runtimes. Unlike plaintext inference, where latency is dominated by FLOPs, in PI non-linear functions (namely ReLU) are the bottleneck. Thus, practical PI demands novel ReLU-aware optimizations. To reduce PI latency we propose a gradient-based algorithm that selectively linearizes ReLUs while maintaining prediction accuracy. We evaluate our algorithm on several standard PI benchmarks. The results demonstrate up to more accuracy (iso-ReLU count at 50K) or less latency (iso-accuracy at 70%) than the current state of the art and advance the Pareto frontier across the latency-accuracy space. To complement empirical results, we present a"no free lunch"theorem that sheds light on how and when network linearization is possible while maintaining prediction accuracy. Public code is available at https://github.com/NYU-DICE-Lab/selective_network_linearization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- LinGCN: Structural Linearized Graph Convolutional Network for Homomorphically Encrypted InferenceHongwu Peng, Ran Ran, Yukui Luo, Jiahui Zhao 等NeurIPS 2023 · 被引用 57 次
- AutoReP: Automatic ReLU Replacement for Fast Private Network InferenceHongwu Peng, Shaoyi Huang, Tong Zhou, Yukui Luo 等ICCV 2023 · 被引用 44 次
- MPCViT: Searching for Accurate and Efficient MPC-Friendly Vision Transformer with Heterogeneous AttentionWenxuan Zeng, Meng Li, Wenjie Xiong, Tong Tong 等ICCV 2023 · 被引用 38 次
- SAL-ViT: Towards Latency Efficient Private Inference on ViT using Selective Attention Search with a Learnable Softmax ApproximationYuke Zhang, Dake Chen, Souvik Kundu, Chenghao Li 等ICCV 2023 · 被引用 30 次
- PrivCirNet: Efficient Private Inference via Block Circulant TransformationTianshi Xu, Lemeng Wu, Runsheng Wang, Meng LiNeurIPS 2024 · 被引用 21 次
它引用的顶会 Paper6
- Oblivious Neural Network Predictions via MiniONN TransformationsJian Liu, Mika Juuti, Yao Lu, N. AsokanCCS 2017 · 被引用 800 次
- CrypTFlow2: Practical 2-Party Secure InferenceDeevashwer Rathee, Mayank Rathee, Nishant Kumar, Nishanth Chandran 等CCS 2020 · 被引用 294 次
- DeepReDuce: ReLU Reduction for Fast Private InferenceNandan Kumar Jha, Zahra Ghodsi, Siddharth Garg, Brandon ReagenICML 2021 · 被引用 108 次
- CryptoNAS: Private Inference on a ReLU BudgetZahra Ghodsi, Akshaj Kumar Veldanda, Brandon Reagen, Siddharth GargNeurIPS 2020 · 被引用 103 次
- SAFENet: A Secure, Accurate and Fast Neural Network InferenceQian Lou, Yilin Shen, Hongxia Jin, Lei JiangICLR 2021 · 被引用 65 次
相关 Paper
- Circa: Stochastic ReLUs for Private Deep LearningZahra Ghodsi, Nandan Kumar Jha, Brandon Reagen, Siddharth GargNeurIPS 2021 · 被引用 41 次
- Learning to Linearize Deep Neural Networks for Secure and Efficient Private InferenceSouvik Kundu, Shunlin Lu, Yuke Zhang, Jacqueline Tiffany Liu 等ICLR 2023 · 被引用 3 次
- Disparate Impact on Group Accuracy of Linearization for Private InferenceSaswat Das, Marco Romanelli, Ferdinando FiorettoICML 2024 · 被引用 4 次
- ReLUPruner: Rethinking ReLU Importance with Taylor Expansion for Efficient Private InferenceZhenpeng Li, Jinshuo Liu, Xinyan Wang, Lina Wang 等AAAI 2026
- Characterizing and Optimizing End-to-End Systems for Private InferenceKarthik Garimella, Zahra Ghodsi, Nandan Kumar Jha, Siddharth Garg 等ASPLOS 2023 · 被引用 15 次
