Lune

ISCA2026顶会

Unlocking Pipeline Parallelism for Bootstrapping: A Pipelined Multi-Chiplet TFHE Accelerator

Yibo Du, Mengdi Wang, Cangyuan Li, Yinhe Han, Ying Wang

2026年份

摘要

TFHE (FHE over the Torus) is a promising fully homomorphic encryption scheme, but it suffers from significant performance overhead. Its performance is severely constrained by the nn iterations of Homomorphic MUXs (HMUXs) within the bootstrapping procedure (n(n is an encryption parameter). To address this, prior works have proposed TFHE accelerators, which process these nn iterations of HMUXs sequentially, fundamentally bottlenecking the overall throughput. In this paper, we explore the previously overlooked pipeline parallelism across HMUXs, overcoming the sequential execution bottleneck to achieve extraordinary throughput. However, unleashing this pipeline parallelism incurs extreme memory bandwidth pressure, as each of the HMUXs requires access to a bootstrapping key (BSK) represented as high-degree polynomials, and nn HMUXs access different BSKs. Conventional HBM-based solutions struggle to handle the massive concurrent BSK accesses needed to sustain the across-HMUX pipeline parallelism. To address this challenge, we propose a multi-chiplet pipelined architecture, which features a distributed SRAM hierarchy to buffer all BSKs so that BSK movement over off-chip memory can be eliminated. By keeping the BSKs resident within distributed on-chip SRAMs, which we call the BSK-distributed strategy, this distributed memory hierarchy confines intensive BSK access within each chiplet, resolving simultaneous BSK access conflicts. Furthermore, this architecture enables both intra- and inter-chiplet polynomial coefficientgrained pipeline to maximize resource utilization. Another challenge arises from the inter-HMUX ciphertext transfers, which could cause frequent die-to-die communication, and this interdie communication has higher latency than intra-die access. To address this, we propose an Interleaved-Fusion policy that fuses multiple contiguous HMUXs into groups and interleaves these groups across chiplets. As the optimal interleaved-fusion mapping varies under different encryption parameters (e.g., different iteration count n)n), we propose an Offline InterleavedFusion Scheduler with an Interleaved-Fusion Cost Model and a dynamic programming algorithm, minimizing the total execution time. Evaluation demonstrates a 3.1×−30.5×\mathbf{3. 1} \times \mathbf{- 3 0. 5} \times performance-perarea improvement on TFHE applications over state-of-the-art TFHE accelerators.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 73377d42-615e-4eb6-bc1b-7c17447c0488

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖