Seesaw: Compensating for Nonlinear Reduction with Linear Computations for Private Inference
Fabing Li, Yuanhao Zhai, Shuangyu Cai, Mingyu Gao
摘要
With increasingly serious data privacy concerns and strict regulations, privacy-preserving machine learning (PPML) has emerged to securely execute machine learning tasks without violating privacy. Unfortunately, the computational cost to securely execute nonlinear computations in PPML remains significant, calling for new model architecture designs with fewer nonlinear operations. We propose Seesaw, a novel neural architecture search method tailored for PPML. Seesaw exploits a previously unexplored opportunity to leverage more linear computations and nonlinear result reuse, in order to compensate for the accuracy loss due to nonlinear reduction. It incorporates specifically designed pruning and search strategies, not only to efficiently handle the much larger design space of both linear and nonlinear operators, but also to achieve a better balance between the model accuracy and the online/offline execution latencies. Compared to the state-of-the-art design for image classification on ImageNet, Seesaw achieves 1.68× lower online latency and 1.55× lower total online + offline latency at 71% iso-accuracy, or 3.65% higher accuracy at iso-latency of 190 seconds, while using much simpler and faster search and training methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM InferenceWenxuan Zeng, Ye Dong, Jinjin Zhou, Jin Tan 等NeurIPS 2025 · 被引用 4 次
- An Efficient Private GPT Never Autoregressively DecodesZhengyi Li, Yue Guan, Kang Yang, Yu Feng 等ICML 2025
它引用的顶会 Paper13
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 被引用 2,107 次
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 被引用 1,075 次
- Oblivious Neural Network Predictions via MiniONN TransformationsJian Liu, Mika Juuti, Yao Lu, N. AsokanCCS 2017 · 被引用 800 次
- DeepReDuce: ReLU Reduction for Fast Private InferenceNandan Kumar Jha, Zahra Ghodsi, Siddharth Garg, Brandon ReagenICML 2021 · 被引用 108 次
- CryptoNAS: Private Inference on a ReLU BudgetZahra Ghodsi, Akshaj Kumar Veldanda, Brandon Reagen, Siddharth GargNeurIPS 2020 · 被引用 103 次
相关 Paper
- PASNet: Polynomial Architecture Search Framework for Two-party Computation-based Secure Neural Network DeploymentHongwu Peng, Shanglin Zhou, Yukui Luo, Nuo Xu 等DAC 2023 · 被引用 5 次
- PP-Stream: Toward High-Performance Privacy-Preserving Neural Network Inference via Distributed Stream ProcessingQingxiu Liu, Qun Huang, Xiang Chen, Sa Wang 等ICDE 2024 · 被引用 8 次
- MD-ML: Super Fast Privacy-Preserving Machine Learning for Malicious Security with a Dishonest MajorityBoshi Yuan, Shixuan Yang, Yongxiang Zhang, Ning Ding 等USENIX Security 2024 · 被引用 22 次
- Selective Network Linearization for Efficient Private InferenceMinsu Cho, Ameya Joshi, Brandon Reagen, Siddharth Garg 等ICML 2022 · 被引用 55 次
- United We Stand: Accelerating Privacy-Preserving Neural Inference by Conjunctive Optimization with Interleaved NexusQiao Zhang, Tao Xiang, Chunsheng Xin, Hongyi WuAAAI 2024
