Morphling: A Throughput-Maximized TFHE-based Accelerator using Transform-domain Reuse
Prasetiyo, Adiwena Putra, Joo-Young Kim
Abstract
Fully Homomorphic Encryption (FHE) has become an increasingly important aspect in modern computing, particularly in preserving privacy in cloud computing by enabling computation directly on encrypted data. Despite its potential, FHE generally poses major computational challenges, including huge computational and memory requirements. The bootstrapping operation, which is essential particularly in Torus-Fhe(tfhe) scheme, involves intensive computations characterized by an enormous number of polynomial multiplications. For instance, performing a single bootstrapping at the 128-bit security level requires more than 10,000 polynomial multiplications. Our in-depth analysis reveals that domain-transform operations, i.e., Fast Fourier Transform (FFT), contribute up to 88% of these operations, which is the bottleneck of the TFHE system. To address these challenges, we propose Morphling, an accelerator architecture that combines the 2D systolic array and strategic use of transform-domain reuse in order to reduce the overhead of domain-transform in TFHE. This novel approach effectively reduces the number of required domain-transform operations by up to 83.3 %, allowing more computational cores in a given die area. In addition, we optimize its micro architecture design for end-to-end TFHE operation, such as merge-split pipelined-FFT for efficient domain-transform operation, double-pointer method for high-throughput polynomial rotation, and specialized buffer design. Furthermore, we introduce custom instructions for tiling, batching, and scheduling of multiple ciphertext operations. This facilitates software-hardware co-optimization, effectively mapping high-level applications such as XG-Boost classifier, Neural-Network, and VGG-9. As a result, Morphling, with four 2D systolic arrays and four vector units with domain-transform reuse, takes 74.79 mm2die area and 53.00 W power consumption in 28nm process. It achieves a throughput of up to 147,615 bootstrappings per second, demonstrating improvements of 3440x over the CPU, 143x over the GPU, and 14.7x over the state-of-the-art TFHE accelerator. It can run various deep learning models with sub-second latency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 44b88420-9785-4f2e-8119-655d2dfc60e9Cited by top-tier papers3
- Trinity: A General Purpose FHE AcceleratorXianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian et al.MICRO 2024 · 34 citations
- ABC-FHE: A Resource-Efficient Accelerator Enabling Bootstrappable Parameters for Client-Side Fully Homomorphic EncryptionSungwoong Yune, Hyojeong Lee, Adiwena Putra, Hyunjun Cho et al.DAC 2025 · 2 citations
- CROPHE: Cross-Operator Dataflow Optimization for Fully Homomorphic Encryption AcceleratorsXinhua Chen, Jiangbin Dong, Hongren Zheng, Tian Tang et al.HPCA 2026 · 1 citation
Related papers
- Strix: An End-to-End Streaming Architecture with Two-Level Ciphertext Batching for Fully Homomorphic Encryption with Programmable BootstrappingAdiwena Putra, Prasetiyo, Yi Chen, John Kim et al.MICRO 2023 · 27 citations
- FlashTFHE: A Scalable Architecture for Efficient Multi-Bit Fully Homomorphic EncryptionJiaao Ma, Ceyu Xu, Ning Liang, Lisa Wu WillsISCA 2026
- FPT: A Fixed-Point Accelerator for Torus Fully Homomorphic EncryptionMichiel Van Beirendonck, Jan-Pieter D'Anvers, Furkan Turan, Ingrid VerbauwhedeCCS 2023 · 28 citations
- MATCHA: a fast and energy-efficient accelerator for fully homomorphic encryption over the torusLei Jiang, Qian Lou, Nrushad JoshiDAC 2022 · 58 citations
- Unlocking Pipeline Parallelism for Bootstrapping: A Pipelined Multi-Chiplet TFHE AcceleratorYibo Du, Mengdi Wang, Cangyuan Li, Yinhe Han et al.ISCA 2026
