Context-Aware Membership Inference Attacks against Pre-trained Large Language Models
Hongyan Chang, Ali Shahin Shamsabadi, Kleomenis Katevas, Hamed Haddadi, Reza Shokri
Abstract
Membership Inference Attacks (MIAs) on pretrained Large Language Models (LLMs) aim at determining if a data point was part of the model's training set. Prior MIAs that are built for classification models fail at LLMs, due to ignoring the generative nature of LLMs across token sequences. In this paper, we present a novel attack on pre-trained LLMs that adapts MIA statistical tests to the perplexity dynamics of subsequences within a data point. Our method significantly outperforms prior approaches, revealing context-dependent memorization patterns in pre-trained LLMs. * The work was completed during Hongyan's internship at Brave.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Exploring the limits of strong membership inference attacks on large language modelsJamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2025 · 26 citations
- Membership Inference Attacks Against Fine-tuned Diffusion Language ModelsYuetian Chen, Kaiyuan Zhang, Yuntao Du, Edoardo Stoppa et al.ICLR 2026 · 6 citations
- Natural Identifiers for Privacy and Data Audits in Large Language ModelsLorenzo Rossi, Bartlomiej Marek, Franziska Boenisch, Adam DziedzicICLR 2026 · 3 citations
- Window-based Membership Inference Attacks Against Fine-tuned Large Language ModelsYuetian Chen, Yuntao Du, Kaiyuan Zhang, Ashish Kundu et al.USENIX Security 2026 · 1 citation
- Privacy Attacks on Image AutoRegressive ModelsAntoni Kowalczuk, Jan Dubinski, Franziska Boenisch, Adam DziedzicICML 2025
Builds on13
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
Related papers
- Towards Label-Only Membership Inference Attack against Pre-trained Large Language ModelsYu He, Boheng Li, Liu Liu, Zhongjie Ba et al.USENIX Security 2025
- ReCaLL: Membership Inference via Relative Conditional Log-LikelihoodsRoy Xie, Junlin Wang, Ruomin Huang, Minxing Zhang et al.EMNLP 2024 · 8 citations
- Membership Inference Attacks on Tokenizers of Large Language ModelsMeng Tong, Yuntao Du, Kejiang Chen, Weiming Zhang et al.USENIX Security 2026
- The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic TextMatthieu Meeus, Lukas Wutschitz, Santiago Zanella-Béguelin, Shruti Tople et al.ICML 2025
- LOMIA: Label-Only Membership Inference Attacks against Pre-trained Large Vision-Language ModelsYihao Liu, Xinqi Lyu, Dong Wang, Yanjie Li et al.NeurIPS 2025 · 3 citations
