Powerful Training-Free Membership Inference Against Fine-Tuned Autoregressive Language Models
David Ilic, David Stanojevic, Kostadin Cvejoski
Abstract
Fine-tuned language models pose significant privacy risks, as they may memorize and expose sensitive information from their training data. Membership inference attacks (MIAs) provide a principled framework for auditing these risks, yet existing methods achieve limited detection rates, particularly at the low falsepositive thresholds required for practical privacy auditing. We present EZ-MIA, a membership inference attack that exploits a key observation: memorization manifests most strongly at error positions, specifically tokens where the model predicts incorrectly yet still shows elevated probability for training examples. We introduce the Error Zone (EZ) score, which measures the directional imbalance of probability shifts at error positions relative to a pretrained reference model. This principled statistic requires only two forward passes per query and no model training of any kind. On WikiText with GPT-2, EZ-MIA achieves 3.8× higher detection than the previous state-of-the-art under identical conditions (66.3% versus 17.5% true positive rate at 1% false positive rate), with near-perfect discrimination (AUC 0.98). At the stringent 0.1% FPR threshold critical for realworld auditing, we achieve 8× higher detection than prior work (14.0% versus 1.8%), requiring no reference model training. These gains extend to larger architectures: on AG News with Llama-2-7B, we achieve 3× higher detection (46.7% versus 15.8% TPR at 1% FPR). These results establish that privacy risks of finetuned language models are substantially greater than previously understood, with implications for both privacy auditing and deployment decisions. Code is available at https://github. com/JetBrains-Research/ez-mia .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe1f7136-a1b6-4119-beff-c247f67e3d0fBuilds on5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
Related papers
- In-Context Probing for Membership Inference in Fine-Tuned Language ModelsZhexi Lu, Hongliang Chi, Nathalie Baracaldo, Swanand Ravindra Kadhe et al.NDSS 2026 · 3 citations
- DF-MIA: A Distribution-Free Membership Inference Attack on Fine-Tuned Large Language ModelsZhiheng Huang, Yannan Liu, Daojing He, Yu LiAAAI 2025 · 7 citations
- Information-Theoretic Membership Inference for Granular Quantification of MemorizationJiashu Tao, Reza ShokriICLR 2026
- Exploring the limits of strong membership inference attacks on large language modelsJamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2025 · 26 citations
- Window-based Membership Inference Attacks Against Fine-tuned Large Language ModelsYuetian Chen, Yuntao Du, Kaiyuan Zhang, Ashish Kundu et al.USENIX Security 2026 · 1 citation
