Hyena: Balancing Packing, Reuse, and Rotations for Encrypted Inference
Sarabjeet Singh, Shreyas Singh, Sumanth Gudaparthi, Xiong Fan, Rajeev Balasubramonian
摘要
Deep neural networks are widely used in a range of commercial services. Many of these services are hosted on the cloud, requiring users to send their personal data to the cloud. This, in turn, exposes the user’s private and sensitive data to several third parties. To address this problem, Homomorphic Encryption (HE) has been introduced, where the user encrypts their data before sending it to the cloud; the cloud performs operations on encrypted data and returns a ciphertext that the user must then decrypt. While this approach keeps user data private, it demands orders of magnitude more computation and data movement. It is, therefore, imperative to design hardware/software techniques to lower the overheads when executing AI services under Homomorphic Encryption schemes.In this paper, we consider a range of HE implementations for AI inference and address the key bottlenecks in state-of-the-art frameworks. We start by making the case for a hybrid HE and Multi-Party Computation (MPC) scheme that is more practical than pure Fully HE. This paper introduces new techniques at various levels: (i) we introduce new data packing techniques that result in lower data movement, (ii) we introduce new dataflows that increase reuse and reduce other costly HE operations (rotations, key switching, NTT conversion), (iii) we evaluate Hyena on a balanced pipelined architecture that efficiently handles the above primitives. The resulting framework, Hyena (new packing + dataflow), achieves better performance and energy than several packing baselines. Compared to the widely used Channel-packing, Hyena is 38× faster and achieves 162× lower energy consumption, with an overall ResNet20 inference end-to-end latency of 11.4 ms, using a 163 mm2 accelerator dissipating 16.75 W.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCTianshi Xu, Wen-jie Lu, Jiangrui Yu, Yi Chen 等USENIX Security 2025
- Fenc2: Unifying Data Packing for Efficient Private Inference via Convolution and Architecture-Aware Fragment EncodingRan Ran, Zhaoting Gong, Nuo Xu, Yuanchao Xu 等ISCA 2026
它引用的顶会 Paper23
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 被引用 1,075 次
- ABY3: A Mixed Protocol Framework for Machine LearningPayman Mohassel, Peter RindalCCS 2018 · 被引用 898 次
- Oblivious Neural Network Predictions via MiniONN TransformationsJian Liu, Mika Juuti, Yao Lu, N. AsokanCCS 2017 · 被引用 800 次
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella 等HPCA 2020 · 被引用 490 次
相关 Paper
- FxHENN: FPGA-based acceleration framework for homomorphic encrypted CNN inferenceYilan Zhu, Xinyao Wang, Lei Ju, Shanqing GuoHPCA 2023 · 被引用 39 次
- Falcon: Fast Spectral Inference on Encrypted DataQian Lou, Wen-jie Lu, Cheng Hong, Lei JiangNeurIPS 2020 · 被引用 50 次
- SpENCNN: Orchestrating Encoding and Sparsity for Fast Homomorphically Encrypted Neural Network InferenceRan Ran, Xinwei Luo, Wei Wang, Tao Liu 等ICML 2023 · 被引用 17 次
- Orion: A Fully Homomorphic Encryption Framework for Deep LearningAustin Ebel, Karthik Garimella, Brandon ReagenASPLOS 2025 · 被引用 40 次
- Cheetah: Optimizing and Accelerating Homomorphic Encryption for Private InferenceBrandon Reagen, Wooseok Choi, Yeongil Ko, Vincent T. Lee 等HPCA 2021 · 被引用 147 次
