NeurIPS2024
PrivAuditor: Benchmarking Data Protection Vulnerabilities in LLM Adaptation Techniques
Derui Zhu, Dingfan Chen, Xiongfei Wu, Jiahui Geng, Zhuo Li, Jens Grossklags, Lei Ma
摘要
Large Language Models (LLMs) are recognized for their potential to be an important building block toward achieving artificial general intelligence due to their unprecedented capability for solving diverse tasks. Despite these achievements, LLMs often underperform in domain-specific tasks without training on relevant domain data. This phenomenon, which is often attributed to distribution shifts, makes adapting pre-trained LLMs with domain-specific data crucial. However, this adaptation raises significant privacy concerns, especially when the data involved come from sensitive domains. In this work, we extensively investigate the privacy vulnerabilities of adapted (fine-tuned) LLMs and benchmark privacy leakage across a wide range of data modalities, state-of-the-art privacy attack methods, adaptation techniques, and model architectures. We systematically evaluate and pinpoint critical factors related to privacy leakage. With our organized codebase and actionable insights, we aim to provide a standardized auditing tool for practitioners seeking to deploy customized LLM applications with faithful privacy assessments. * Equal contribution 38th Conference on Neural Information Processing Systems (NeurIPS 2024) Track on Datasets and Benchmarks. including model size and the degree of training data repetition, have been presented [10, 15, 16, 17 ]. Yet, in the context of fine-tuning/adaptation scenarios, recent privacy risk assessments have typically been limited to specific model architectures (mainly encoder-based models), a narrow selection of fine-tuning methods, and a certain choice of attack methods [7, 8, 9, 10, 11, 18] . A comprehensive benchmark evaluation is still missing, despite its importance for providing critical insights and accurate privacy assessments to facilitate the practical application of domain-specific LLMs. In particular, this gap highlights a crucial research question: To what extent, and in what ways, do different adaptation methods influence the privacy risk of LLMs? To address the research question, this paper presents, to the best of our knowledge, the first benchmark investigating the privacy implications of LLM adaptation techniques, accompanied by a comprehensive empirical study. We focus on membership inference attack (MIA) techniques [19] , which aim to determine whether a given query sample was used for adapting the target LLM, due to their popularity and close relationship to a broader class of topics [12, 20, 21] . Our investigation encompasses five types of LLMs with different architectures (T5 [3], LLaMA [22] , OPT [23], BLOOM [24] , and GPT-J [25]), seven LLM adaptation techniques representative of the current state of the art, and three datasets from different domains that closely mimic real-world sensitive fields. With our presented benchmark and comprehensive study, we aim to provide critical insights into the privacy risks associated with LLM adaptation techniques and guide the secure development of new models. Privacy Measurement for Large Language Models We evaluate the privacy vulnerabilities of LLMs through the lens of MIAs [19] , which are widely recognized for their extensive applicability. MIAs are also closely associated with other privacy concerns, such as training data reconstruction [12, 15] and the retrieval of personally identifiable information [13, 26, 14] , underscoring its critical role in privacy assessments.
