A Unified Perspective on Adversarial Membership Manipulation in Vision Models
RUIZE GAO, Kaiwen Zhou, Yongqiang Chen, Feng Liu
摘要
Membership inference attacks (MIAs) aim to determine whether a specific data point was part of a model’s training set, serving as effective tools for evaluating privacy leakage of vision models. However, existing MIAs implicitly assume honest query inputs, and their adversarial robustness remains unexplored. We show that MIAs for vision models expose a previously overlooked adversarial surface: adversarial membership manipulation, where imperceptible perturbations can reliably push non-member images into the “member’’ region of state-of-the-art MIAs. In this paper, we provide the first unified perspective on this phenomenon by analyzing its mechanism and implications. We begin by demonstrating that adversarial membership fabrication is consistently effective across diverse architectures and datasets. We then reveal a distinctive geometric signature—a characteristic gradient-norm collapse trajectory—that reliably separates fabricated from true members despite their nearly identical semantic representations. Building on this insight, we introduce a principled detection strategy grounded in gradient-geometry signals and develop a robust inference framework that substantially mitigates adversarial manipulation. Extensive experiments show that fabrication is broadly effective, while our detection and robust inference strategies significantly enhance resilience. This work establishes the first comprehensive framework for adversarial membership manipulation in vision models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song 等S&P 2022 · 被引用 1,049 次
- Label-Only Membership Inference AttacksChristopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, Nicolas PapernotICML 2021 · 被引用 628 次
相关 Paper
- Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference AttackJing Xue, Zhishen Sun, Haishan Ye, Luo Luo 等AAAI 2026
- Mixup Training for Generative Models to Defend Membership Inference AttacksZhe Ji, Qiansiqi Hu, Liyao Xiang, Chenghu ZhouINFOCOM 2023 · 被引用 3 次
- Leveraging Adversarial Examples to Quantify Membership Information LeakageGanesh Del Grosso, Hamid Jalalzai, Georg Pichler, Catuscia Palamidessi 等CVPR 2022 · 被引用 2 次
- Membership Inference Attacks against Vision Transformers: Mosaic MixUp Training to the DefenseQiankun Zhang, Di Yuan, Boyu Zhang, Bin Yuan 等CCS 2024 · 被引用 1 次
- Membership Inference Attacks by Exploiting Loss TrajectoryYiyong Liu, Zhengyu Zhao, Michael Backes, Yang ZhangCCS 2022 · 被引用 79 次
