Lune

MICRO2023Top-tier venue

Strix: An End-to-End Streaming Architecture with Two-Level Ciphertext Batching for Fully Homomorphic Encryption with Programmable Bootstrapping

Adiwena Putra, Prasetiyo, Yi Chen, John Kim, Joo-Young Kim

2023Year
27Citations
4Top-tier citations

Abstract

Homomorphic encryption (HE) is a type of cryptography that allows computations to be performed on encrypted data. The technique relies on learning with errors problem, where data is hidden under noise for security. To avoid accumulating too much noise, the process of bootstrapping is needed to reset the noise level in the ciphertext, but it requires a large bootstrapping key and is computationally expensive. The fully homomorphic encryption over the torus (TFHE) scheme offers a faster and programmable bootstrapping (PBS) algorithm, which is crucial for security-focused applications like machine learning. Nonetheless, the current TFHE scheme does not support ciphertext packing, resulting in low-throughput performance. To the best of our knowledge, this is the first work that thoroughly analyzes TFHE bootstrapping, identifies the TFHE acceleration bottleneck in GPUs, and proposes a hardware TFHE accelerator to solve the bottleneck.

We begin by identifying the TFHE acceleration bottleneck in GPUs, which is caused by the blind rotation fragmentation problem. This issue can be significantly improved by increasing the batch size in PBS. We propose a two-level batching approach to substantially enhance the batch size in PBS. In order to efficiently implement this solution, we propose Strix, which utilizes streaming and fully pipelined architecture with specialized functional units to accelerate the sequential ciphertext processing in TFHE. In particular, we propose a novel microarchitecture for decomposition in TFHE, suitable for processing streaming data at high throughput. We also utilize a fully-pipelined FFT microarchitecture to eliminate the complex memory access bottleneck and enhance its performance through a folding scheme, resulting in 2× throughput improvement and 1.7× area reduction. Strix achieves over 1, 067× and 37× higher throughput in running TFHE with PBS than the state-of-the-art implementation on CPU and GPU, respectively, outperforming the state of the art TFHE accelerator by 7.4×.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 0a207c3c-b350-45c5-adc0-957ad0457bcf

Cited by top-tier papers4

Ask how each one uses it

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines