Composing Differential Privacy and Secure Computation: A Case Study on Scaling Private Record Linkage
Xi He, Ashwin Machanavajjhala, Cheryl J. Flynn, Divesh Srivastava
摘要
Private record linkage (PRL) is the problem of identifying pairs of records that are similar as per an input matching rule from databases held by two parties that do not trust one another. We identify three key desiderata that a PRL solution must ensure: (1) perfect precision and high recall of matching pairs, (2) a proof of end-to-end privacy, and (3) communication and computational costs that scale subquadratically in the number of input records. We show that all of the existing solutions for PRL? including secure 2-party computation (S2PC), and their variants that use non-private or differentially private (DP) blocking to ensure subquadratic cost -- violate at least one of the three desiderata. In particular, S2PC techniques guarantee end-to-end privacy but have either low recall or quadratic cost. In contrast, no end-to-end privacy guarantee has been formalized for solutions that achieve subquadratic cost. This is true even for solutions that compose DP and S2PC: DP does not permit the release of any exact information about the databases, while S2PC algorithms for PRL allow the release of matching records. In light of this deficiency, we propose a novel privacy model, called output constrained differential privacy, that shares the strong privacy protection of DP, but allows for the truthful release of the output of a certain function applied to the data. We apply this to PRL, and show that protocols satisfying this privacy model permit the disclosure of the true matching records, but their execution is insensitive to the presence or absence of a single non-matching record. We find that prior work that combine DP and S2PC techniques even fail to satisfy this end-to-end privacy model. Hence, we develop novel protocols that provably achieve this end-to-end privacy guarantee, together with the other two desiderata of PRL. Our empirical evaluation also shows that our protocols obtain high recall, scale near linearly in the size of the input databases and the output set of matching pairs, and have communication and computational costs that are at least 2 orders of magnitude smaller than S2PC baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen 等VLDB 2020 · 被引用 259 次
- SECRECY: Secure collaborative analytics in untrusted cloudsJohn Liagouris, Vasiliki Kalavri, Muhammad Faisal, Mayank VariaNSDI 2023 · 被引用 53 次
- Falcon: A Privacy-Preserving and Interpretable Vertical Federated Learning SystemYuncheng Wu, Naili Xing, Gang Chen, Tien Tuan Anh Dinh 等VLDB 2023 · 被引用 47 次
- Crypt?: Crypto-Assisted Differential Privacy on Untrusted ServersAmrita Roy Chowdhury, Chenghong Wang, Xi He, Ashwin Machanavajjhala 等SIGMOD 2020 · 被引用 40 次
- Improving Utility and Security of the Shuffler-based Differential PrivacyTianhao Wang, Min Xu, Bolin Ding, Jingren Zhou 等VLDB 2020 · 被引用 39 次
相关 Paper
- SFour: A Protocol for Cryptographically Secure Record Linkage at ScaleBasit Khurram, Florian KerschbaumICDE 2020 · 被引用 11 次
- Cryptographically Secure Private Record Linkage Using Locality-Sensitive HashingRuidi Wei, Florian KerschbaumVLDB 2024 · 被引用 10 次
- Privacy-Preserving Screening for Record LinkageChenyu Huang, Fan Zhang, Huangxun Chen, Yongjun Zhao 等ICDE 2025
- Budget Sharing for Multi-Analyst Differential PrivacyDavid Pujol, Yikai Wu, Brandon Fain, Ashwin MachanavajjhalaVLDB 2021 · 被引用 7 次
- Differentially Private Prototypes for Imbalanced Transfer LearningDariush Wahdany, Matthew Jagielski, Adam Dziedzic, Franziska BoenischAAAI 2025 · 被引用 4 次
