Lune

EMNLP2023Top-tier venue

MoPe: Model Perturbation based Privacy Attacks on Language Models

Marvin Li, Jason Wang, Jeffrey G. Wang, Seth Neel

2023Year
8Citations
10Top-tier citations

Abstract

Recent work has shown that Large Language Models (LLMs) can unintentionally leak sensitive information present in their training data. In this paper, we present MoPe θ (Model Perturbations), a new method to identify with high confidence if a given text is in the training data of a pre-trained language model, given white-box access to the models parameters. MoPe θ adds noise to the model in parameter space and measures the drop in log-likelihood at a given point x, a statistic we show approximates the trace of the Hessian matrix with respect to model parameters. Across language models ranging from 70M to 12B parameters, we show that MoPe θ is more effective than existing loss-based attacks and recently proposed perturbation-based methods. We also examine the role of training point order and model size in attack success, and empirically demonstrate that MoPe θ accurately approximate the trace of the Hessian in practice. Our results show that the loss of a point alone is insufficient to determine extractability-there are training points we can recover using our method that have average loss. This casts some doubt on prior works that use the loss of a point as evidence of memorization or "unlearning." * Alphabetical order; equal contribution.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e109f946-e7b2-432f-b99e-60458ecc8743

Cited by top-tier papers10

Ask how each one uses it

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines