Cinnamon: A Framework for Scale-Out Encrypted AI
Siddharth Jayashankar, Edward Chen, Tom Tang, Wenting Zheng, Dimitrios Skarlatos
Abstract
Fully homomorphic encryption (FHE) is a promising cryptographic solution that enables computation on encrypted data, but its adoption remains a challenge due to steep performance overheads. Although recent FHE architectures have made valiant efforts to narrow the performance gap, they not only have massive monolithic chip designs but also only target small ML workloads. We present Cinnamon, a framework for accelerating state-of-the-art ML workloads that are encrypted using FHE. Cinnamon accelerates encrypted computing by exploiting parallelism at all levels of a program, using novel algorithms, compilers, and hardware techniques to create a scale-out design for FHE as opposed to a monolithic chip design. Our evaluation of the Cinnamon framework on small programs shows a 2.3× improvement in performance compared to prior state-of-the-art designs. Further, we use Cinnamon to show for the first time the scalability of large ML models such as the BERT language model in FHE. Cinnamon achieves a speedup of 36,600× compared to a CPU bringing down the inference time from 17 hours to 1.67 seconds thereby enabling new opportunities for privacy-preserving machine learning. Finally, Cinnamon's parallelization strategies and architectural extensions reduce the required resources per-chip leading to a 5× and 2.68× improvement in performance-per-dollar compared to state-of-the-art monolithic and chiplet architectures respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Bridging Usability and Performance: A Tensor Compiler for Autovectorizing Homomorphic EncryptionEdward Chen, Fraser Brown, Wenting ZhengUSENIX Security 2026 · 3 citations
- Leveraging ASIC AI Chips for Homomorphic EncryptionJianming Tong, Tianhao Huang, Jingtian Dang, Leo de Castro et al.HPCA 2026 · 2 citations
- Fenc2: Unifying Data Packing for Efficient Private Inference via Convolution and Architecture-Aware Fragment EncodingRan Ran, Zhaoting Gong, Nuo Xu, Yuanchao Xu et al.ISCA 2026
Builds on19
- Spectre Attacks: Exploiting Speculative ExecutionPaul Kocher, Jann Horn, Anders Fogh, Daniel Genkin et al.S&P 2019 · 2,435 citations
- Meltdown: Reading Kernel Memory from User SpaceMoritz Lipp, Michael Schwarz, Daniel Gruss, Thomas Prescher et al.USENIX Security 2018 · 1,456 citations
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- F1: A Fast and Programmable Accelerator for Fully Homomorphic EncryptionNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Srinivas Devadas et al.MICRO 2021 · 294 citations
- CraterLake: a hardware accelerator for efficient unbounded computation on encrypted dataNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Nathan Manohar et al.ISCA 2022 · 205 citations
Related papers
- FxHENN: FPGA-based acceleration framework for homomorphic encrypted CNN inferenceYilan Zhu, Xinyao Wang, Lei Ju, Shanqing GuoHPCA 2023 · 39 citations
- EFFACT: A Highly Efficient Full-Stack FHE Acceleration PlatformYi Huang, Xinsheng Gong, Xiangyu Kong, Dibei Chen et al.HPCA 2025 · 10 citations
- FHE-CGRA: Enable Efficient Acceleration of Fully Homomorphic Encryption on CGRAsMiaomiao Jiang, Yilan Zhu, Honghui You, Cheng Tan et al.DAC 2024 · 5 citations
- Tricycle: Private Transformer Inference with Tricyclic EncodingsLawrence Lim, Vikas Kalagi, Julia Novick, Jiaming Liu et al.CCS 2026 · 10 citations
- Cheetah: Optimizing and Accelerating Homomorphic Encryption for Private InferenceBrandon Reagen, Wooseok Choi, Yeongil Ko, Vincent T. Lee et al.HPCA 2021 · 147 citations
