PP-Stream: Toward High-Performance Privacy-Preserving Neural Network Inference via Distributed Stream Processing
Qingxiu Liu, Qun Huang, Xiang Chen, Sa Wang, Wenhao Wang, Shujie Han, Patrick P. C. Lee
Abstract
Privacy preservation is critical for neural network inference, which often involves collaborative execution of different parties to make predictions on sensitive data based on sensitive neural network models. However, the expensive cryptographic operations of privacy preservation also pose performance chal-lenges to neural network inference. We address this performance-security tension by designing PP-Stream, a distributed stream processing system for high-performance privacy-preserving neural network inference. PP-Stream adopts hybrid privacy-preserving mechanisms for linear and non-linear operations of neural network inference. It treats inference data as real-time data streams, and parallelizes the inference operations across multiple pipelined stages that are executed by multiple servers and threads. It also solves the load-balanced resource allocation across servers and threads as an optimization problem. We prototype PP-Stream and show via testbed experiments that it achieves low inference latencies on various neural network models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2fd8971-e80c-4c7d-a10f-e4cd05472d6aCited by top-tier papers2
- Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference ServingShihong Gao, Xin Zhang, Yanyan Shen, Lei ChenSIGMOD 2025 · 7 citations
- On the (In-)Security of the Shuffling Defense in the Transformer Secure InferenceZhengyi Li, Yakai Wang, Jingwen Leng, Kang Yang et al.ACL 2026
Builds on14
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- BatchCrypt: Efficient Homomorphic Encryption for Cross-Silo Federated LearningChengliang Zhang, Suyi Li, Junzhe Xia, Wei Wang et al.USENIX ATC 2020 · 967 citations
- ABY3: A Mixed Protocol Framework for Machine LearningPayman Mohassel, Peter RindalCCS 2018 · 898 citations
- Oblivious Neural Network Predictions via MiniONN TransformationsJian Liu, Mika Juuti, Yao Lu, N. AsokanCCS 2017 · 800 citations
Related papers
- Delphi: A Cryptographic Inference Service for Neural NetworksPratyush Mishra, Ryan Lehmkuhl, Akshayaram Srinivasan, Wenting Zheng et al.USENIX Security 2020
- PPGNN: Fast and Accurate Privacy-Preserving Graph Neural Network Inference via Parallel and Pipelined Arithmetic-and-Logic FHE AcceleratorYuntao Wei, Xueyan Wang, Song Bian, Yicheng Huang et al.DAC 2024 · 5 citations
- CrypTFlow2: Practical 2-Party Secure InferenceDeevashwer Rathee, Mayank Rathee, Nishant Kumar, Nishanth Chandran et al.CCS 2020 · 294 citations
- HEMET: A Homomorphic-Encryption-Friendly Privacy-Preserving Mobile Neural Network ArchitectureQian Lou, Lei JiangICML 2021 · 88 citations
- Characterizing and Optimizing End-to-End Systems for Private InferenceKarthik Garimella, Zahra Ghodsi, Nandan Kumar Jha, Siddharth Garg et al.ASPLOS 2023 · 15 citations
