DeepReDuce: ReLU Reduction for Fast Private Inference
Nandan Kumar Jha, Zahra Ghodsi, Siddharth Garg, Brandon Reagen
摘要
The recent rise of privacy concerns has led researchers to devise methods for private neural inference -- where inferences are made directly on encrypted data, never seeing inputs. The primary challenge facing private inference is that computing on encrypted data levies an impractically-high latency penalty, stemming mostly from non-linear operators like ReLU. Enabling practical and private inference requires new optimization methods that minimize network ReLU counts while preserving accuracy. This paper proposes DeepReDuce: a set of optimizations for the judicious removal of ReLUs to reduce private inference latency. The key insight is that not all ReLUs contribute equally to accuracy. We leverage this insight to drop, or remove, ReLUs from classic networks to significantly reduce inference latency and maintain high accuracy. Given a target network, DeepReDuce outputs a Pareto frontier of networks that tradeoff the number of ReLUs and accuracy. Compared to the state-of-the-art for private inference DeepReDuce improves accuracy and reduces ReLU count by up to 3.5% (iso-ReLU count) and 3.5 (iso-accuracy), respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing 等NeurIPS 2022 · 被引用 209 次
- LinGCN: Structural Linearized Graph Convolutional Network for Homomorphically Encrypted InferenceHongwu Peng, Ran Ran, Yukui Luo, Jiahui Zhao 等NeurIPS 2023 · 被引用 57 次
- Selective Network Linearization for Efficient Private InferenceMinsu Cho, Ameya Joshi, Brandon Reagen, Siddharth Garg 等ICML 2022 · 被引用 55 次
- CryptoGCN: Fast and Scalable Homomorphically Encrypted Graph Convolutional Network InferenceRan Ran, Wei Wang, Quan Gang, Jieming Yin 等NeurIPS 2022 · 被引用 52 次
- AutoReP: Automatic ReLU Replacement for Fast Private Network InferenceHongwu Peng, Shaoyi Huang, Tong Zhou, Yukui Luo 等ICCV 2023 · 被引用 44 次
它引用的顶会 Paper8
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 被引用 2,107 次
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 被引用 1,075 次
- Oblivious Neural Network Predictions via MiniONN TransformationsJian Liu, Mika Juuti, Yao Lu, N. AsokanCCS 2017 · 被引用 800 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- CryptoNAS: Private Inference on a ReLU BudgetZahra Ghodsi, Akshaj Kumar Veldanda, Brandon Reagen, Siddharth GargNeurIPS 2020 · 被引用 103 次
相关 Paper
- Learning to Linearize Deep Neural Networks for Secure and Efficient Private InferenceSouvik Kundu, Shunlin Lu, Yuke Zhang, Jacqueline Tiffany Liu 等ICLR 2023 · 被引用 3 次
- Circa: Stochastic ReLUs for Private Deep LearningZahra Ghodsi, Nandan Kumar Jha, Brandon Reagen, Siddharth GargNeurIPS 2021 · 被引用 41 次
- ReLUPruner: Rethinking ReLU Importance with Taylor Expansion for Efficient Private InferenceZhenpeng Li, Jinshuo Liu, Xinyan Wang, Lina Wang 等AAAI 2026
- Disparate Impact on Group Accuracy of Linearization for Private InferenceSaswat Das, Marco Romanelli, Ferdinando FiorettoICML 2024 · 被引用 4 次
- PAPER: Privacy-Preserving Convolutional Neural Networks using Low-Degree Polynomial Approximations and Structural Optimizations on Leveled FHEEduardo Chielle, Manaar Alam, Jinting Liu, Jovan Kascelan 等CCS 2026
