Optimization and Robustness-Informed Membership Inference Attacks for LLMs
Zichen Song, Qixin Zhang, Ming Li, Yao Shu
Abstract
The proliferation of Large Language Models (LLMs) has raised concerns over training data privacy. Membership Inference Attacks (MIA), aiming to identify whether specific data was used for training, pose significant privacy risks. However, existing MIA methods struggle to address the scale and complexity of modern LLMs. This paper introduces OR-MIA, a novel MIA framework inspired by model optimization and input robustness. First, training data points are expected to exhibit smaller gradient norms due to optimization dynamics. Second, member samples show greater stability, with gradient norms being less sensitive to controlled input perturbations. OR-MIA leverages these principles by perturbing inputs, computing gradient norms, and using them as features for a robust classifier to distinguish members from non-members. Evaluations on LLMs (70M to 6B parameters) and various datasets demonstrate that OR-MIA outperforms existing methods, achieving over 90% accuracy. Our findings highlight a critical vulnerability in LLMs and underscore the need for improved privacy-preserving training paradigms. Optimization and Robustness-Informed Membership Inference Attacks for LLMs Quantum Machine Learning: A Review of Algorithms and Applications Artificial Intelligence is revolutionizing the way we approach healthcare data analytics. This paper explores the applications of deep learning in predicting patient outcomes. This paper belongs to the field of quantum computing and machine learning and is classified as: Quantum Computing
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 619b78eb-a090-473a-809a-bae9c14f6e25Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
Related papers
- Order of Magnitude Speedups for LLM Membership InferenceRongting Zhang, Martin Bertran Lopez, Aaron RothEMNLP 2024 · 1 citation
- LOMIA: Label-Only Membership Inference Attacks against Pre-trained Large Vision-Language ModelsYihao Liu, Xinqi Lyu, Dong Wang, Yanjie Li et al.NeurIPS 2025 · 3 citations
- Membership Inference Attacks against Large Vision-Language ModelsZhan Li, Yongtao Wu, Yihang Chen, Francesco Tonin et al.NeurIPS 2024 · 43 citations
- Robust Membership Inference for Large Language Models under Adversarial Generative CorruptionYuanhong Huang, Huili Wang, Xueying Bai, Jinrui Wang et al.ACL 2026
- Decoding Web Memorization: A Semantic Membership Inference Attack on LLMsZhiyao Wu, Zi Liang, Haibo HuWWW 2026
