Can We Infer Confidential Properties of Training Data from LLMs?
Pengrun Huang, Chhavi Yadav, Kamalika Chaudhuri, Ruihan Wu
Abstract
Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets to support applications in fields such as healthcare, finance, and law. These fine-tuning datasets often have sensitive and confidential dataset-level propertiessuch as patient demographics or disease prevalence-that are not intended to be revealed. While prior work has studied property inference attacks on discriminative models (e.g., image classification models) and generative models (e.g., GANs for image data), it remains unclear if such attacks transfer to LLMs. In this work, we introduce PropInfer, a benchmark task for evaluating property inference in LLMs under two fine-tuning paradigms: question-answering and chat-completion. Built on the ChatDoctor dataset, our benchmark includes a range of property types and task configurations. We further propose two tailored attacks: a prompt-based generation attack and a shadow-model attack leveraging word frequency signals. Empirical evaluations across multiple pretrained LLMs show the success of our attacks, revealing a previously unrecognized vulnerability in LLMs. We release our code at github.com/PengrunH/Property_inference_attack_LLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d399f41-4993-4d3c-bfb6-62e9da730c09Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
Related papers
- Can Personal Health Information Be Secured in LLM? Privacy Attack and Defense in the Medical DomainYujin Kang, Eunsun Kim, Yoon-Sik ChoCCS 2025
- Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question AnsweringHwan Chang, Yumin Kim, Yonghyun Jun, Hwanhee LeeEMNLP 2025
- Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMsEyal German, Sagiv Antebi, Daniel Samira, Asaf Shabtai et al.ICLR 2026 · 3 citations
- Quantifying Privacy Risks of Prompts in Visual Prompt LearningYixin Wu, Rui Wen, Michael Backes, Pascal Berrang et al.USENIX Security 2024 · 12 citations
- A Benchmark for Semantic Sensitive Information in LLMs OutputsQingjie Zhang, Han Qiu, Di Wang, Yiming Li et al.ICLR 2025
