Lune

ISCA2026Top-tier venue

Unlocking Pipeline Parallelism for Bootstrapping: A Pipelined Multi-Chiplet TFHE Accelerator

Yibo Du, Mengdi Wang, Cangyuan Li, Yinhe Han, Ying Wang

2026Year

Abstract

TFHE (FHE over the Torus) is a promising fully homomorphic encryption scheme, but it suffers from significant performance overhead. Its performance is severely constrained by the nn iterations of Homomorphic MUXs (HMUXs) within the bootstrapping procedure (n(n is an encryption parameter). To address this, prior works have proposed TFHE accelerators, which process these nn iterations of HMUXs sequentially, fundamentally bottlenecking the overall throughput. In this paper, we explore the previously overlooked pipeline parallelism across HMUXs, overcoming the sequential execution bottleneck to achieve extraordinary throughput. However, unleashing this pipeline parallelism incurs extreme memory bandwidth pressure, as each of the HMUXs requires access to a bootstrapping key (BSK) represented as high-degree polynomials, and nn HMUXs access different BSKs. Conventional HBM-based solutions struggle to handle the massive concurrent BSK accesses needed to sustain the across-HMUX pipeline parallelism. To address this challenge, we propose a multi-chiplet pipelined architecture, which features a distributed SRAM hierarchy to buffer all BSKs so that BSK movement over off-chip memory can be eliminated. By keeping the BSKs resident within distributed on-chip SRAMs, which we call the BSK-distributed strategy, this distributed memory hierarchy confines intensive BSK access within each chiplet, resolving simultaneous BSK access conflicts. Furthermore, this architecture enables both intra- and inter-chiplet polynomial coefficientgrained pipeline to maximize resource utilization. Another challenge arises from the inter-HMUX ciphertext transfers, which could cause frequent die-to-die communication, and this interdie communication has higher latency than intra-die access. To address this, we propose an Interleaved-Fusion policy that fuses multiple contiguous HMUXs into groups and interleaves these groups across chiplets. As the optimal interleaved-fusion mapping varies under different encryption parameters (e.g., different iteration count n)n), we propose an Offline InterleavedFusion Scheduler with an Interleaved-Fusion Cost Model and a dynamic programming algorithm, minimizing the total execution time. Evaluation demonstrates a 3.1×−30.5×\mathbf{3. 1} \times \mathbf{- 3 0. 5} \times performance-perarea improvement on TFHE applications over state-of-the-art TFHE accelerators.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 73377d42-615e-4eb6-bc1b-7c17447c0488

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines