Lune

MICRO2023顶会

Strix: An End-to-End Streaming Architecture with Two-Level Ciphertext Batching for Fully Homomorphic Encryption with Programmable Bootstrapping

Adiwena Putra, Prasetiyo, Yi Chen, John Kim, Joo-Young Kim

2023年份
27被引次数
4顶会引用

摘要

Homomorphic encryption (HE) is a type of cryptography that allows computations to be performed on encrypted data. The technique relies on learning with errors problem, where data is hidden under noise for security. To avoid accumulating too much noise, the process of bootstrapping is needed to reset the noise level in the ciphertext, but it requires a large bootstrapping key and is computationally expensive. The fully homomorphic encryption over the torus (TFHE) scheme offers a faster and programmable bootstrapping (PBS) algorithm, which is crucial for security-focused applications like machine learning. Nonetheless, the current TFHE scheme does not support ciphertext packing, resulting in low-throughput performance. To the best of our knowledge, this is the first work that thoroughly analyzes TFHE bootstrapping, identifies the TFHE acceleration bottleneck in GPUs, and proposes a hardware TFHE accelerator to solve the bottleneck.

We begin by identifying the TFHE acceleration bottleneck in GPUs, which is caused by the blind rotation fragmentation problem. This issue can be significantly improved by increasing the batch size in PBS. We propose a two-level batching approach to substantially enhance the batch size in PBS. In order to efficiently implement this solution, we propose Strix, which utilizes streaming and fully pipelined architecture with specialized functional units to accelerate the sequential ciphertext processing in TFHE. In particular, we propose a novel microarchitecture for decomposition in TFHE, suitable for processing streaming data at high throughput. We also utilize a fully-pipelined FFT microarchitecture to eliminate the complex memory access bottleneck and enhance its performance through a folding scheme, resulting in 2× throughput improvement and 1.7× area reduction. Strix achieves over 1, 067× and 37× higher throughput in running TFHE with PBS than the state-of-the-art implementation on CPU and GPU, respectively, outperforming the state of the art TFHE accelerator by 7.4×.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖