Morphling: A Throughput-Maximized TFHE-based Accelerator using Transform-domain Reuse
Prasetiyo, Adiwena Putra, Joo-Young Kim
摘要
Fully Homomorphic Encryption (FHE) has become an increasingly important aspect in modern computing, particularly in preserving privacy in cloud computing by enabling computation directly on encrypted data. Despite its potential, FHE generally poses major computational challenges, including huge computational and memory requirements. The bootstrapping operation, which is essential particularly in Torus-Fhe(tfhe) scheme, involves intensive computations characterized by an enormous number of polynomial multiplications. For instance, performing a single bootstrapping at the 128-bit security level requires more than 10,000 polynomial multiplications. Our in-depth analysis reveals that domain-transform operations, i.e., Fast Fourier Transform (FFT), contribute up to 88% of these operations, which is the bottleneck of the TFHE system. To address these challenges, we propose Morphling, an accelerator architecture that combines the 2D systolic array and strategic use of transform-domain reuse in order to reduce the overhead of domain-transform in TFHE. This novel approach effectively reduces the number of required domain-transform operations by up to 83.3 %, allowing more computational cores in a given die area. In addition, we optimize its micro architecture design for end-to-end TFHE operation, such as merge-split pipelined-FFT for efficient domain-transform operation, double-pointer method for high-throughput polynomial rotation, and specialized buffer design. Furthermore, we introduce custom instructions for tiling, batching, and scheduling of multiple ciphertext operations. This facilitates software-hardware co-optimization, effectively mapping high-level applications such as XG-Boost classifier, Neural-Network, and VGG-9. As a result, Morphling, with four 2D systolic arrays and four vector units with domain-transform reuse, takes 74.79 mm2die area and 53.00 W power consumption in 28nm process. It achieves a throughput of up to 147,615 bootstrappings per second, demonstrating improvements of 3440x over the CPU, 143x over the GPU, and 14.7x over the state-of-the-art TFHE accelerator. It can run various deep learning models with sub-second latency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Trinity: A General Purpose FHE AcceleratorXianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian 等MICRO 2024 · 被引用 34 次
- ABC-FHE: A Resource-Efficient Accelerator Enabling Bootstrappable Parameters for Client-Side Fully Homomorphic EncryptionSungwoong Yune, Hyojeong Lee, Adiwena Putra, Hyunjun Cho 等DAC 2025 · 被引用 2 次
- CROPHE: Cross-Operator Dataflow Optimization for Fully Homomorphic Encryption AcceleratorsXinhua Chen, Jiangbin Dong, Hongren Zheng, Tian Tang 等HPCA 2026 · 被引用 1 次
相关 Paper
- Strix: An End-to-End Streaming Architecture with Two-Level Ciphertext Batching for Fully Homomorphic Encryption with Programmable BootstrappingAdiwena Putra, Prasetiyo, Yi Chen, John Kim 等MICRO 2023 · 被引用 27 次
- FlashTFHE: A Scalable Architecture for Efficient Multi-Bit Fully Homomorphic EncryptionJiaao Ma, Ceyu Xu, Ning Liang, Lisa Wu WillsISCA 2026
- FPT: A Fixed-Point Accelerator for Torus Fully Homomorphic EncryptionMichiel Van Beirendonck, Jan-Pieter D'Anvers, Furkan Turan, Ingrid VerbauwhedeCCS 2023 · 被引用 28 次
- MATCHA: a fast and energy-efficient accelerator for fully homomorphic encryption over the torusLei Jiang, Qian Lou, Nrushad JoshiDAC 2022 · 被引用 58 次
- Unlocking Pipeline Parallelism for Bootstrapping: A Pipelined Multi-Chiplet TFHE AcceleratorYibo Du, Mengdi Wang, Cangyuan Li, Yinhe Han 等ISCA 2026
