Prism: Private Relational Data Synthesis with Language Models
Guohui Guan, Chang Ge
摘要
Organizations increasingly rely on relational data containing sensitive personal information for downstream tasks. A common approach is to release synthetic data as a privacy-preserving substitute, typically by learning a generative model and sampling from it. Differential privacy (DP) provides a formal guarantee of data privacy, and existing data synthesis systems aim to balance its inherent trade-off between data privacy and task utility. The advent of pretrained large language models (LLMs) has reshaped this trade-off landscape: on the utility side, LLMs possess stronger representational capabilities and encode extensive prior knowledge that does not consume the privacy budget, leading to improved task utility; on the other hand, enforcing differential privacy for LLM-based relational data synthesis introduces new technical challenges. We present P rism, an end-to-end framework for DP relational data synthesis using LLMs. P rism privatizes knowledge transfer from an ensemble of teacher LLMs to a student generator via DP output aggregation. We introduce three key innovations in the aggregation mechanism to address LLM-specific challenges, including 1) token pruning with tries to minimize teacher queries and hence the privacy cost; 2) a data-dependent DP mechanism that adaptively scales noise for efficient budget use; and 3) probabilistic sampling with threshold filtering to reduce bias and preserve diversity. Extensive empirical results across twelve experiments show that Prism achieves substantially higher predictive utility than eight state-of-the-art methods under the same privacy budget.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Contrastive Private Data Synthesis via Weighted Multi-PLM FusionTianyuan Zou, Yang Liu, Peng Li, Yufei Xiong 等ICML 2025
- Privacy Preserving In-Context-Learning Framework for Large Language ModelsBishnu Bhusal, Manoj Acharya, Ramneet Kaur, Colin Samplawski 等AAAI 2026 · 被引用 1 次
- Differentially Private Preference Data Synthesis for Large Language Model AlignmentFengyu Gao, Jing YangICML 2026
- PEARL: Differentially Private and Entropy-Aware Regulated Language GenerationSeongho Joo, Hyukhun Koh, Kyomin JungICML 2026
- PrivCode: When Code Generation Meets Differential PrivacyZheng Liu, Chen Gong, Terry Yue Zhuo, Kecen Li 等NDSS 2026 · 被引用 5 次
