USENIX Security2025Top-tier venue
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
Yu He, Boheng Li, Liu Liu, Zhongjie Ba, Wei Dong, Yiming Li, Zhan Qin, Kui Ren, Chun Chen
Abstract
Membership Inference Attacks (MIAs) aim to predict whether a data sample belongs to the model's training set or not. Although prior research has extensively explored MIAs in Large Language Models (LLMs), they typically require accessing to complete output logits (, logits-based attacks), which are usually not available in practice. In this paper, we study the vulnerability of pre-trained LLMs to MIAs in the label-only setting, where the adversary can only access generated tokens (text). We first reveal that existing label-only MIAs have minor effects in attacking pre-trained LLMs, although they are highly effective in inferring fine-tuning datasets used for personalized LLMs. We find that their failure stems from two main reasons, including better generalization and overly coarse perturbation. Specifically, due to the extensive pre-training corpora and exposing each sample only a few times, LLMs exhibit minimal robustness differences between members and non-members. This makes token-level perturbations too coarse to capture such differences. To alleviate these problems, we propose PETAL: a label-only membership inference attack based on PEr-Token semAntic simiLarity. Specifically, PETAL leverages token-level semantic similarity to approximate output probabilities and subsequently calculate the perplexity. It finally exposes membership based on the common assumption that members are `better' memorized and have smaller perplexity. We conduct extensive experiments on the WikiMIA benchmark and the more challenging MIMIR benchmark. Empirically, our PETAL performs better than the extensions of existing label-only attacks against personalized LLMs and even on par with other advanced logit-based attacks across all metrics on five prevalent open-source LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5bf3410-78bb-4969-ac42-dddc2b6f33d1Cited by top-tier papers13
- Context-Aware Membership Inference Attacks against Pre-trained Large Language ModelsHongyan Chang, Ali Shahin Shamsabadi, Kleomenis Katevas, Hamed Haddadi et al.EMNLP 2025 · 19 citations
- Membership Inference Attacks Against Fine-tuned Diffusion Language ModelsYuetian Chen, Kaiyuan Zhang, Yuntao Du, Edoardo Stoppa et al.ICLR 2026 · 6 citations
- Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!Zhexin Zhang, Yuhao Sun, Junxiao Yang, Shiyao Cui et al.ICLR 2026 · 5 citations
- Toward Efficient Membership Inference Attacks Against Federated Large Language Models: A Projection Residual ApproachGuilin Deng, Silong Chen, Yuchuan Luo, Yi Liu et al.S&P 2026 · 4 citations
- LOMIA: Label-Only Membership Inference Attacks against Pre-trained Large Vision-Language ModelsYihao Liu, Xinqi Lyu, Dong Wang, Yanjie Li et al.NeurIPS 2025 · 3 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- Membership Inference Attacks on Tokenizers of Large Language ModelsMeng Tong, Yuntao Du, Kejiang Chen, Weiming Zhang et al.USENIX Security 2026
- Exploring the limits of strong membership inference attacks on large language modelsJamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2025 · 26 citations
- ReCaLL: Membership Inference via Relative Conditional Log-LikelihoodsRoy Xie, Junlin Wang, Ruomin Huang, Minxing Zhang et al.EMNLP 2024 · 8 citations
- Decoding Web Memorization: A Semantic Membership Inference Attack on LLMsZhiyao Wu, Zi Liang, Haibo HuWWW 2026
- Order of Magnitude Speedups for LLM Membership InferenceRongting Zhang, Martin Bertran Lopez, Aaron RothEMNLP 2024 · 1 citation
