Neural Network Inversion in Adversarial Setting via Background Knowledge Alignment
Ziqi Yang, Jiyi Zhang, Ee-Chien Chang, Zhenkai Liang
摘要
The wide application of deep learning technique has raised new security concerns about the training data and test data. In this work, we investigate the model inversion problem under adversarial settings, where the adversary aims at inferring information about the target model's training data and test data from the model's prediction values. We develop a solution to train a second neural network that acts as the inverse of the target model to perform the inversion. The inversion model can be trained with black-box accesses to the target model. We propose two main techniques towards training the inversion model in the adversarial settings. First, we leverage the adversary's background knowledge to compose an auxiliary set to train the inversion model, which does not require access to the original training data. Second, we design a truncation-based technique to align the inversion model to enable effective inversion of the target model from partial predictions that the adversary obtains on victim user's data. We systematically evaluate our approach in various machine learning tasks and model architectures on multiple image datasets. We also confirm our results on Amazon Rekognition, a commercial prediction API that offers "machine learning as a service". We show that even with partial knowledge about the black-box model's training data, and with only partial prediction values, our inversion approach is still able to perform accurate inversion of the target model, and outperform previous approaches.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper57
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 被引用 395 次
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- ProPILE: Probing Privacy Leakage in Large Language ModelsSiwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri 等NeurIPS 2023 · 被引用 229 次
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir 等NeurIPS 2022 · 被引用 196 次
- R-GAP: Recursive Gradient Attack on PrivacyJunyi Zhu, Matthew B. BlaschkoICLR 2021 · 被引用 157 次
相关 Paper
- Reinforcement Learning-Based Black-Box Model Inversion AttacksGyojin Han, Jaehyun Choi, Haeil Lee, Junmo KimCVPR 2023
- Are Your Sensitive Attributes Private? Novel Model Inversion Attribute Inference Attacks on Classification ModelsShagufta Mehnaz, Sayanton V. Dibbo, Ehsanul Kabir, Ninghui Li 等USENIX Security 2022
- Adaptive Domain Inference Attack with Concept HierarchyYuechun Gu, Jiajie He, Keke ChenKDD 2025
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural NetworksYuheng Zhang, Ruoxi Jia, Hengzhi Pei, Wenxiao Wang 等CVPR 2020
