Unlocking Pipeline Parallelism for Bootstrapping: A Pipelined Multi-Chiplet TFHE Accelerator
Yibo Du, Mengdi Wang, Cangyuan Li, Yinhe Han, Ying Wang
Abstract
TFHE (FHE over the Torus) is a promising fully homomorphic encryption scheme, but it suffers from significant performance overhead. Its performance is severely constrained by the iterations of Homomorphic MUXs (HMUXs) within the bootstrapping procedure is an encryption parameter). To address this, prior works have proposed TFHE accelerators, which process these iterations of HMUXs sequentially, fundamentally bottlenecking the overall throughput. In this paper, we explore the previously overlooked pipeline parallelism across HMUXs, overcoming the sequential execution bottleneck to achieve extraordinary throughput. However, unleashing this pipeline parallelism incurs extreme memory bandwidth pressure, as each of the HMUXs requires access to a bootstrapping key (BSK) represented as high-degree polynomials, and HMUXs access different BSKs. Conventional HBM-based solutions struggle to handle the massive concurrent BSK accesses needed to sustain the across-HMUX pipeline parallelism. To address this challenge, we propose a multi-chiplet pipelined architecture, which features a distributed SRAM hierarchy to buffer all BSKs so that BSK movement over off-chip memory can be eliminated. By keeping the BSKs resident within distributed on-chip SRAMs, which we call the BSK-distributed strategy, this distributed memory hierarchy confines intensive BSK access within each chiplet, resolving simultaneous BSK access conflicts. Furthermore, this architecture enables both intra- and inter-chiplet polynomial coefficientgrained pipeline to maximize resource utilization. Another challenge arises from the inter-HMUX ciphertext transfers, which could cause frequent die-to-die communication, and this interdie communication has higher latency than intra-die access. To address this, we propose an Interleaved-Fusion policy that fuses multiple contiguous HMUXs into groups and interleaves these groups across chiplets. As the optimal interleaved-fusion mapping varies under different encryption parameters (e.g., different iteration count , we propose an Offline InterleavedFusion Scheduler with an Interleaved-Fusion Cost Model and a dynamic programming algorithm, minimizing the total execution time. Evaluation demonstrates a performance-perarea improvement on TFHE applications over state-of-the-art TFHE accelerators.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 73377d42-615e-4eb6-bc1b-7c17447c0488Related papers
- Strix: An End-to-End Streaming Architecture with Two-Level Ciphertext Batching for Fully Homomorphic Encryption with Programmable BootstrappingAdiwena Putra, Prasetiyo, Yi Chen, John Kim et al.MICRO 2023 · 27 citations
- FlashTFHE: A Scalable Architecture for Efficient Multi-Bit Fully Homomorphic EncryptionJiaao Ma, Ceyu Xu, Ning Liang, Lisa Wu WillsISCA 2026
- Affinity-based Optimizations for TFHE on Processing-in-DRAMKevin Nam, Heon Hui Jung, Hyunyoung Oh, Yunheung PaekASPLOS 2025 · 2 citations
- Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-DesignZhiwei Wang, Haoqi He, Lutan Zhao, Qingyun Niu et al.ASPLOS 2026
- BTS: an accelerator for bootstrappable fully homomorphic encryptionSangpyo Kim, Jongmin Kim, Michael Jaemin Kim, Wonkyung Jung et al.ISCA 2022 · 184 citations
