SpENCNN: Orchestrating Encoding and Sparsity for Fast Homomorphically Encrypted Neural Network Inference
Ran Ran, Xinwei Luo, Wei Wang, Tao Liu, Gang Quan, Xiaolin Xu, Caiwen Ding, Wujie Wen
摘要
Homomorphic Encryption (HE) is a promising technology to protect clients' data privacy for Machine Learning as a Service (MLaaS) on public clouds. However, HE operations can be orders of magnitude slower than their counterparts for plaintexts and thus result in prohibitively high inference latency, seriously hindering the practicality of HE. In this paper, we propose a HE-based fast neural network (NN) inference framework-SpENCNN built upon the co-design of HE operation-aware model sparsity and the single-instruction-multiple-data (SIMD)-friendly data packing, to improve NN inference latency. In particular, we first develop an encryption-aware HE-group convolution technique that can partition channels among different groups based on the data size and ciphertext size, and then encode them into the same ciphertext by novel groupinterleaved encoding, so as to dramatically reduce the number of bottlenecked operations in HE convolution. We further tailor a HE-friendly sub-block weight pruning to reduce the costly HE-based convolution operation. Our experiments show that SpENCNN can achieve overall speedups of 8.37×, 12.11×, 19.26×, and 1.87× for LeNet, VGG-5, HEFNet, and ResNet-20 respectively, with negligible accuracy loss. Our code is publicly available at https://github. com/ranran0523/SPECNN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- PrivCirNet: Efficient Private Inference via Block Circulant TransformationTianshi Xu, Lemeng Wu, Runsheng Wang, Meng LiNeurIPS 2024 · 被引用 21 次
- Testing and Understanding Deviation Behaviors in FHE-Hardened Machine Learning ModelsYiteng Peng, Daoyuan Wu, Zhibo Liu, Dongwei Xiao 等ICSE 2025 · 被引用 1 次
- Fenc2: Unifying Data Packing for Efficient Private Inference via Convolution and Architecture-Aware Fragment EncodingRan Ran, Zhaoting Gong, Nuo Xu, Yuanchao Xu 等ISCA 2026
- An Efficient Private GPT Never Autoregressively DecodesZhengyi Li, Yue Guan, Kang Yang, Yu Feng 等ICML 2025
- Secure Transformer Inference Made Non-interactiveJiawen Zhang, Xinpeng Yang, Lipeng He, Kejia Chen 等NDSS 2025
它引用的顶会 Paper7
- Enabling Deep Spiking Neural Networks with Hybrid Conversion and Spike Timing Dependent BackpropagationNitin Rathi, Gopalakrishnan Srinivasan, Priyadarshini Panda, Kaushik RoyICLR 2020 · 被引用 347 次
- CraterLake: a hardware accelerator for efficient unbounded computation on encrypted dataNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Nathan Manohar 等ISCA 2022 · 被引用 205 次
- Low-Complexity Deep Convolutional Neural Networks on Fully Homomorphic Encryption Using Multiplexed Parallel ConvolutionsEunsang Lee, Joon-Woo Lee, Junghyun Lee, Young-Sik Kim 等ICML 2022 · 被引用 171 次
- DeepReDuce: ReLU Reduction for Fast Private InferenceNandan Kumar Jha, Zahra Ghodsi, Siddharth Garg, Brandon ReagenICML 2021 · 被引用 108 次
- CryptoNAS: Private Inference on a ReLU BudgetZahra Ghodsi, Akshaj Kumar Veldanda, Brandon Reagen, Siddharth GargNeurIPS 2020 · 被引用 103 次
相关 Paper
- FxHENN: FPGA-based acceleration framework for homomorphic encrypted CNN inferenceYilan Zhu, Xinyao Wang, Lei Ju, Shanqing GuoHPCA 2023 · 被引用 39 次
- Falcon: Fast Spectral Inference on Encrypted DataQian Lou, Wen-jie Lu, Cheng Hong, Lei JiangNeurIPS 2020 · 被引用 50 次
- HEMET: A Homomorphic-Encryption-Friendly Privacy-Preserving Mobile Neural Network ArchitectureQian Lou, Lei JiangICML 2021 · 被引用 88 次
- HEPrune: Fast Private Training of Deep Neural Networks With Encrypted Data PruningYancheng Zhang, Mengxin Zheng, Yuzhang Shang, Xun Chen 等NeurIPS 2024 · 被引用 23 次
- FicGCN: Unveiling the Homomorphic Encryption Efficiency from Irregular Graph Convolutional NetworksZhaoxuan Kan, Husheng Han, Shangyi Shi, Tenghui Hua 等ICML 2025
