Fragments to Facts: Partial-Information Fragment Inference from LLMs
Lucas Rosenblatt, Bin Han, Robert Wolfe, Bill Howe
Abstract
Large language models (LLMs) can leak sensitive training data through memorization and membership inference attacks. Prior work has primarily focused on strong adversarial assumptions, including attacker access to entire samples or long, ordered prefixes, leaving open the question of how vulnerable LLMs are when adversaries have only partial, unordered sample information. For example, if an attacker knows a patient has "hypertension," under what conditions can they query a model fine-tuned on patient data to learn the patient also has "osteoarthritis?" In this paper, we introduce a more general threat model under this weaker assumption and show that fine-tuned LLMs are susceptible to these fragment-specific extraction attacks. To systematically investigate these attacks, we propose two data-blind methods: (1) a likelihood ratio attack inspired by methods from membership inference, and (2) a novel approach, PRISM, which regularizes the ratio by leveraging an external prior. Using examples from medical and legal settings, we show that both methods are competitive with a data-aware baseline classifier that assumes access to labeled indistribution data, underscoring their robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36469729-b158-4e28-b2bb-51367fdadd23Builds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
Related papers
- Quantifying Privacy Risks of Masked Language Models Using Membership Inference AttacksFatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick et al.EMNLP 2022 · 72 citations
- Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMsEyal German, Sagiv Antebi, Daniel Samira, Asaf Shabtai et al.ICLR 2026 · 3 citations
- Retracing the Past: LLMs Emit Training Data When They Get LostMyeongseob Ko, Nikhil Reddy Billa, Adam Nguyen, Charles Fleming et al.EMNLP 2025 · 1 citation
- Decoding Web Memorization: A Semantic Membership Inference Attack on LLMsZhiyao Wu, Zi Liang, Haibo HuWWW 2026
- Teach LLMs to Phish: Stealing Private Information from Language ModelsAshwinee Panda, Christopher A. Choquette-Choo, Zhengming Zhang, Yaoqing Yang et al.ICLR 2024 · 41 citations
