Differentially Private Set Union
Sivakanth Gopi, Pankaj Gulhane, Janardhan Kulkarni, Judy Hanwen Shen, Milad Shokouhi, Sergey Yekhanin
摘要
We study the basic operation of set union in the global model of differential privacy. In this problem, we are given a universe of items, possibly of infinite size, and a database of users. Each user contributes a subset of items. We want an (,)-differentially private algorithm which outputs a subset such that the size of is as large as possible. The problem arises in countless real world applications; it is particularly ubiquitous in natural language processing (NLP) applications as vocabulary extraction. For example, discovering words, sentences, -grams etc., from private text data belonging to users is an instance of the set union problem.Known algorithms for this problem proceed by collecting a subset of items from each user, taking the union of such subsets, and disclosing the items whose noisy counts fall above a certain threshold. Crucially, in the above process, the contribution of each individual user is always independent of the items held by other users, resulting in a wasteful aggregation process, where some item counts happen to be way above the threshold. We deviate from the above paradigm by allowing users to contribute their items in a dependent fashion, guided by a policy. In this new setting ensuring privacy is significantly delicate. We prove that any policy which has certain contractive properties would result in a differentially private algorithm. We design two new algorithms for differentially private set union, one using Laplace noise and other Gaussian noise, which use -contractive and -contractive policies respectively and provide concrete examples of such policies. Our experiments show that the new algorithms in combination with our policies significantly outperform previously known mechanisms for the problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Differentially Private Synthetic Data via Foundation Model APIs 1: ImagesZinan Lin, Sivakanth Gopi, Janardhan Kulkarni, Harsha Nori 等ICLR 2024 · 被引用 63 次
- Differentially Private n-gram ExtractionKunho Kim, Sivakanth Gopi, Janardhan Kulkarni, Sergey YekhaninNeurIPS 2021 · 被引用 22 次
- Incorporating Item Frequency for Differentially Private Set UnionRicardo Silva Carvalho, Ke Wang, Lovedeep Singh GondaraAAAI 2022 · 被引用 12 次
- Sparsity-Preserving Differentially Private Training of Large Embedding ModelsBadih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar 等NeurIPS 2023 · 被引用 9 次
- Counting Distinct Elements Under Person-Level Differential PrivacyThomas Steinke, Alexander KnopNeurIPS 2023 · 被引用 4 次
它引用的顶会 Paper3
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- BLENDER: Enabling Local Search with a Hybrid Differential Privacy ModelBrendan Avent, Aleksandra Korolova, David Zeber, Torgeir Hovden 等USENIX Security 2017 · 被引用 101 次
- Differentially Private n-gram ExtractionKunho Kim, Sivakanth Gopi, Janardhan Kulkarni, Sergey YekhaninNeurIPS 2021 · 被引用 22 次
相关 Paper
- Scalable Private Partition Selection via Adaptive WeightingJustin Y. Chen, Vincent Cohen-Addad, Alessandro Epasto, Morteza ZadimoghaddamICML 2025
- Differentially Private Domain DiscoveryVinod Raman, Travis Dick, Matthew JosephICLR 2026
- Private Set Union with Multiple ContributionsTravis Dick, Haim Kaplan, Alex Kulesza, Uri Stemmer 等NeurIPS 2025
- Algorithms for bounding contribution for histogram estimation under user-level privacyYuhan Liu, Ananda Theertha Suresh, Wennan Zhu, Peter Kairouz 等ICML 2023 · 被引用 14 次
- Smoothly Bounding User Contributions in Differential PrivacyAlessandro Epasto, Mohammad Mahdian, Jieming Mao, Vahab S. Mirrokni 等NeurIPS 2020 · 被引用 17 次
