Lune

ICLR2025Top-tier venue

Does Training with Synthetic Data Truly Protect Privacy?

Yunpeng Zhao, Jie Zhang

2025Year
6Top-tier citations

Abstract

As synthetic data becomes increasingly popular in machine learning tasks, numerous methods-without formal differential privacy guarantees-use synthetic data for training. These methods often claim, either explicitly or implicitly, to protect the privacy of the original training data. In this work, we explore four different training paradigms: coreset selection, dataset distillation, data-free knowledge distillation, and synthetic data generated from diffusion models. While all these methods utilize synthetic data for training, they lead to vastly different conclusions regarding privacy preservation. We caution that empirical approaches to preserving data privacy require careful and rigorous evaluation; otherwise, they risk providing a false sense of privacy.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c392f68f-e392-42c6-9354-af340b2c3dbd

Cited by top-tier papers6

Ask how each one uses it

Builds on26

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines