Lune

CVPR2021顶会

On the Difficulty of Membership Inference Attacks

Shahbaz Rezaei, Xin Liu

2021年份
17顶会引用

摘要

Recent studies propose membership inference (MI) attacks on deep models, where the goal is to infer if a sample has been used in the training process. Despite their apparent success, these studies only report accuracy, precision, and recall of the positive class (member class). Hence, the performance of these attacks have not been clearly reported on negative class (non-member class). In this paper, we show that the way the MI attack performance has been reported is often misleading because they suffer from high false positive rate or false alarm rate (FAR) that has not been reported. FAR shows how often the attack model mislabel non-training samples (non-member) as training (member) ones. The high FAR makes MI attacks fundamentally impractical, which is particularly more significant for tasks such as membership inference where the majority of samples in reality belong to the negative (non-training) class. Moreover, we show that the current MI attack models can only identify the membership of misclassified samples with mediocre accuracy at best, which only constitute a very small portion of training samples.

We analyze several new features that have not been comprehensively explored for membership inference before, including distance to the decision boundary and gradient norms, and conclude that deep models' responses are mostly similar among train and non-train samples. We conduct several experiments on image classification tasks, including MNIST, CIFAR-10, CIFAR-100, and Ima-geNet, using various model architecture, including LeNet, AlexNet, ResNet, etc. We show that the current stateof-the-art MI attacks cannot achieve high accuracy and low FAR at the same time, even when the attacker is given several advantages. The source code is available at https://github.com/shrezaei/MI-Attack. Dataset Cifar-100 Cifar-100 Cifar-100 Model AlexNet ResNet DenseNet Target Model Train Acc. 92.48% 95.80% 99.98% Target Model Test Acc. 43.87% 74.14% 82.83% Attack Acc. 82.62% 79.13% 87.74% Attack Precision 91.90% 87.3% 86.97% Attack Recall 86.92% 87.85% 98.29% Attack F1 89.23% 87.45% 92.26% Attack Bal. Acc. 74.02% 61.70% 66.65% Attack FAR 38.89% 64.45% 65.00%

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper17

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖