Lune

USENIX Security2021Top-tier venue

GForce: GPU-Friendly Oblivious and Rapid Neural Network Inference

Lucien K. L. Ng, Sherman S. M. Chow

2021Year
60Citations
14Top-tier citations

Abstract

Neural-network classification is getting more pervasive. It captures data of the subjects to be classified, e.g., appearance for facial recognition, which is personal and often sensitive. Oblivious inference protects the data privacy of both the query and the model. However, it is not as fast and as accurate as its plaintext counterpart. A recent cryptographic solution Delphi (Usenix Security 2020) strives for low latency by using GPU on linear layers and replacing some non-linear units in the model at a price of accuracy. It can handle a query on CIFAR-100 with ∼68% accuracy in 14s or ∼66% accuracy in 2.6s. We propose GForce, tackling the latency issue from the root causes instead of approximating non-linear computations. With the SWALP training approach (ICML 2019), we propose stochastic rounding and truncation (SRT) layers, which fuse quantization with dequantization between non-linear and linear layers and free us from floating-point operations for efficiency. They also ensure high accuracy while working over the severely-finite cryptographic field. We further propose a suite of GPU-friendly secure online/offline protocols for common operations, including comparison and wrap-around handling, which benefit non-linear layers, including our SRT. With our two innovations, GForce supports VGG-16, attaining ∼73% accuracy over CIFAR-100 for the first time, in 0.4s. Compared with the prior best non-approximated solution (Usenix Security 2018), GForce speeds up non-linear layers in VGG by >34×. Our techniques shed light on a new direction that utilizes GPU throughout the model to minimize latency. * Supported by General Research Fund (CUHK 14210319) of UGC, HK. model to the clients for evaluation is often impossible, not to say its financial and privacy implications. Oblivious inference resolves this dilemma. The server with a deep neural network DNN(•) can return the classification result DNN(x) to any client while remains oblivious about x without leaking its model DNN(•). From the perspective of computation nature, a neural network can be divided into linear layers and non-linear layers. Cryptographic solutions often handle linear layers and non-linear layers separately, such as using additive homomorphic encryption (AHE) and garbled circuits (GC), respectively, but these tools impose high overheads. A recurrent research problem is how to perform secure computations of non-linear functions efficiently.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4adb6efe-020e-420e-aae0-0ae925aa3404

Cited by top-tier papers14

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines