MetaKernel: Enabling Efficient Encrypted Neural Network Inference through Unified MVM and Convolution
Peng Yuan, Yan Liu, Jianxin Lai, Long Li, Tianxiang Sui, Linjie Xiao, Xiaojing Zhang, Qing Zhu, Jingling Xue
摘要
Practical encrypted neural network inference under the CKKS fully homomorphic encryption (FHE) scheme relies heavily on accelerating two key kernel operations: Matrix-Vector Multiplication (MVM) and Convolution (Conv). However, existing solutions—such as expert-tuned libraries and domain-specific languages—are designed in an ad hoc manner, leading to significant inefficiencies caused by excessive rotations. We introduce MKR, a novel composition-based compiler approach that optimizes MVM and Conv kernel operations for DNN models under CKKS within a unified framework. MKR decomposes each kernel into composable units, called MetaKernels , to enhance SIMD parallelism within ciphertexts (via horizontal batching) and computational parallelism across them (via vertical batching). Our approach tackles previously unaddressed challenges, including reducing rotation overhead through a rotation-aware cost model for data packing, while also ensuring high slot utilization, uniform handling of inputs with arbitrary sizes, and compatibility with the output tensor layout. Implemented in a production-quality FHE compiler, MKR achieves inference time speedups of 10.08×−185.60× for individual MVM and Conv kernels and 1.75×−11.84× for end-to-end inference compared to a state-of-the-art FHE compiler. Moreover, MKR enables homomorphic execution of large DNN models, where prior methods fail, significantly advancing the practicality of FHE compilers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Orbit: Optimizing Rescale and Bootstrap Placement with Integer Linear Programming Techniques for Secure InferenceZikai Zhou, William Seo, Edward Chen, Alex Ozdemir 等USENIX Security 2026
- FHE-CGRA: Enable Efficient Acceleration of Fully Homomorphic Encryption on CGRAsMiaomiao Jiang, Yilan Zhu, Honghui You, Cheng Tan 等DAC 2024 · 被引用 5 次
- FxHENN: FPGA-based acceleration framework for homomorphic encrypted CNN inferenceYilan Zhu, Xinyao Wang, Lei Ju, Shanqing GuoHPCA 2023 · 被引用 39 次
- BitPacker: Enabling High Arithmetic Efficiency in Fully Homomorphic Encryption AcceleratorsNikola Samardzic, Daniel SánchezASPLOS 2024 · 被引用 19 次
- SpENCNN: Orchestrating Encoding and Sparsity for Fast Homomorphically Encrypted Neural Network InferenceRan Ran, Xinwei Luo, Wei Wang, Tao Liu 等ICML 2023 · 被引用 17 次
