Reinforcement Learning-Based Black-Box Model Inversion Attacks
Gyojin Han, Jaehyun Choi, Haeil Lee, Junmo Kim
Abstract
Model inversion attacks are a type of privacy attack that reconstructs private data used to train a machine learning model, solely by accessing the model. Recently, white-box model inversion attacks leveraging Generative Adversarial Networks (GANs) to distill knowledge from public datasets have been receiving great attention because of their excellent attack performance. On the other hand, current blackbox model inversion attacks that utilize GANs suffer from issues such as being unable to guarantee the completion of the attack process within a predetermined number of query accesses or achieve the same level of performance as whitebox attacks. To overcome these limitations, we propose a reinforcement learning-based black-box model inversion attack. We formulate the latent space search as a Markov Decision Process (MDP) problem and solve it with reinforcement learning. Our method utilizes the confidence scores of the generated images to provide rewards to an agent. Finally, the private data can be reconstructed using the latent vectors found by the agent trained in the MDP. The experiment results on various datasets and models demonstrate that our attack successfully recovers the private information of the target model by achieving state-of-the-art attack performance. We emphasize the importance of studies on privacy-preserving machine learning by proposing a more advanced black-box model inversion attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Label-Only Model Inversion Attacks via Knowledge TransferNgoc-Bao Nguyen, Keshigeyan Chandrasegaran, Milad Abdollahzadeh, Ngai-Man CheungNeurIPS 2023 · 44 citations
- Be Careful What You Smooth For: Label Smoothing Can Be a Privacy Shield but Also a Catalyst for Model Inversion AttacksLukas Struppek, Dominik Hintersdorf, Kristian KerstingICLR 2024 · 26 citations
- Pseudo-Private Data Guided Model Inversion AttacksXiong Peng, Bo Han, Feng Liu, Tongliang Liu et al.NeurIPS 2024 · 12 citations
- Trap-MID: Trapdoor-based Defense against Model Inversion AttacksZhenTing Liu, ShangTse ChenNeurIPS 2024 · 12 citations
- StablePrompt : Automatic Prompt Tuning using Reinforcement Learning for Large Language ModelMinchan Kwon, Gaeun Kim, Jongsuk Kim, Haeil Lee et al.EMNLP 2024 · 11 citations
Builds on8
- Neural Network Inversion in Adversarial Setting via Background Knowledge AlignmentZiqi Yang, Jiyi Zhang, Ee-Chien Chang, Zhenkai LiangCCS 2019 · 257 citations
- Variational Model Inversion AttacksKuan-Chieh Wang, Yan Fu, Ke Li, Ashish Khisti et al.NeurIPS 2021 · 142 citations
- Knowledge-Enriched Distributional Model Inversion AttacksSi Chen, Mostafa Kahla, Ruoxi Jia, Guo-Jun QiICCV 2021 · 124 citations
- Rethinking the Truly Unsupervised Image-to-Image TranslationKyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo et al.ICCV 2021 · 115 citations
- Label-Only Model Inversion Attacks via Boundary RepulsionMostafa Kahla, Si Chen, Hoang Anh Just, Ruoxi JiaCVPR 2022 · 60 citations
Related papers
- Are Your Sensitive Attributes Private? Novel Model Inversion Attribute Inference Attacks on Classification ModelsShagufta Mehnaz, Sayanton V. Dibbo, Ehsanul Kabir, Ninghui Li et al.USENIX Security 2022
- The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural NetworksYuheng Zhang, Ruoxi Jia, Hengzhi Pei, Wenxiao Wang et al.CVPR 2020
- Pseudo Label-Guided Model Inversion Attack via Conditional Generative Adversarial NetworkXiaojian Yuan, Kejiang Chen, Jie Zhang, Weiming Zhang et al.AAAI 2023 · 57 citations
- Face Reconstruction from Facial Templates by Learning Latent Space of a Generator NetworkHatef Otroshi-Shahreza, Sébastien MarcelNeurIPS 2023 · 48 citations
- Query-efficient Attack for Black-box Image Inpainting Forensics via Reinforcement LearningXianbo Mo, Shunquan Tan, Bin Li, Jiwu HuangAAAI 2025 · 5 citations
