DeepReDuce: ReLU Reduction for Fast Private Inference
Nandan Kumar Jha, Zahra Ghodsi, Siddharth Garg, Brandon Reagen
Abstract
The recent rise of privacy concerns has led researchers to devise methods for private neural inference -- where inferences are made directly on encrypted data, never seeing inputs. The primary challenge facing private inference is that computing on encrypted data levies an impractically-high latency penalty, stemming mostly from non-linear operators like ReLU. Enabling practical and private inference requires new optimization methods that minimize network ReLU counts while preserving accuracy. This paper proposes DeepReDuce: a set of optimizations for the judicious removal of ReLUs to reduce private inference latency. The key insight is that not all ReLUs contribute equally to accuracy. We leverage this insight to drop, or remove, ReLUs from classic networks to significantly reduce inference latency and maintain high accuracy. Given a target network, DeepReDuce outputs a Pareto frontier of networks that tradeoff the number of ReLUs and accuracy. Compared to the state-of-the-art for private inference DeepReDuce improves accuracy and reduces ReLU count by up to 3.5% (iso-ReLU count) and 3.5 (iso-accuracy), respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77af5784-ebc5-4d62-94b3-8500272d009aCited by top-tier papers25
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing et al.NeurIPS 2022 · 209 citations
- LinGCN: Structural Linearized Graph Convolutional Network for Homomorphically Encrypted InferenceHongwu Peng, Ran Ran, Yukui Luo, Jiahui Zhao et al.NeurIPS 2023 · 57 citations
- Selective Network Linearization for Efficient Private InferenceMinsu Cho, Ameya Joshi, Brandon Reagen, Siddharth Garg et al.ICML 2022 · 55 citations
- CryptoGCN: Fast and Scalable Homomorphically Encrypted Graph Convolutional Network InferenceRan Ran, Wei Wang, Quan Gang, Jieming Yin et al.NeurIPS 2022 · 52 citations
- AutoReP: Automatic ReLU Replacement for Fast Private Network InferenceHongwu Peng, Shaoyi Huang, Tong Zhou, Yukui Luo et al.ICCV 2023 · 44 citations
Builds on8
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- Oblivious Neural Network Predictions via MiniONN TransformationsJian Liu, Mika Juuti, Yao Lu, N. AsokanCCS 2017 · 800 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- CryptoNAS: Private Inference on a ReLU BudgetZahra Ghodsi, Akshaj Kumar Veldanda, Brandon Reagen, Siddharth GargNeurIPS 2020 · 103 citations
Related papers
- Learning to Linearize Deep Neural Networks for Secure and Efficient Private InferenceSouvik Kundu, Shunlin Lu, Yuke Zhang, Jacqueline Tiffany Liu et al.ICLR 2023 · 3 citations
- Circa: Stochastic ReLUs for Private Deep LearningZahra Ghodsi, Nandan Kumar Jha, Brandon Reagen, Siddharth GargNeurIPS 2021 · 41 citations
- ReLUPruner: Rethinking ReLU Importance with Taylor Expansion for Efficient Private InferenceZhenpeng Li, Jinshuo Liu, Xinyan Wang, Lina Wang et al.AAAI 2026
- Disparate Impact on Group Accuracy of Linearization for Private InferenceSaswat Das, Marco Romanelli, Ferdinando FiorettoICML 2024 · 4 citations
- PAPER: Privacy-Preserving Convolutional Neural Networks using Low-Degree Polynomial Approximations and Structural Optimizations on Leveled FHEEduardo Chielle, Manaar Alam, Jinting Liu, Jovan Kascelan et al.CCS 2026
