Lune

NeurIPS2021顶会

Class-Disentanglement and Applications in Adversarial Detection and Defense

Kaiwen Yang, Tianyi Zhou, Yonggang Zhang, Xinmei Tian, Dacheng Tao

出版方
2021年份
49被引次数
12顶会引用

摘要

What is the minimum necessary information required by a neural net D(•) from an image x to accurately predict its class? Extracting such information in the input space from x can allocate the areas D(•) mainly attending to and shed novel insights to the detection and defense of adversarial attacks. In this paper, we propose "class-disentanglement" that trains a variational autoencoder G(•) to extract this class-dependent information as x -G(x) via a trade-off between reconstructing x by G(x) and classifying x by D(x -G(x)), where the former competes with the latter in decomposing x so the latter retains only necessary information for classification in x -G(x). We apply it to both clean images and their adversarial images and discover that the perturbations generated by adversarial attacks mainly lie in the class-dependent part x -G(x). The decomposition results also provide novel interpretations to classification and attack models. Inspired by these observations, we propose to conduct adversarial detection and adversarial defense respectively on x -G(x) and G(x), which consistently outperform the results on the original x. In experiments, this simple approach substantially improves the detection and defense against different types of adversarial attacks. Code is available: https://github.com/kai-wen-yang/CD-VAE .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper12

问问它们各自怎么用它

它引用的顶会 Paper5

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖