Synthesizing Privacy Preserving Entity Resolution Datasets
Xuedi Qin, Chengliang Chai, Nan Tang, Jian Li, Yuyu Luo, Guoliang Li, Yaoyu Zhu
摘要
Entity resolution (ER) is a core problem in data integration. Many companies have lots of datasets where ER needs to be conducted to integrate the data. On the one hand, it is nontrivial for non-ER experts within companies to design ER solutions. On the other hand, most companies are reluctant to release their real datasets for multiple reasons (e.g., privacy issues). A typical solution from the machine learning (ML) and the statistical community is to create surrogate (a.k.a. analogous) datasets based on the real dataset, release these surrogate datasets to the public to train ML models, such that these models trained on surrogate datasets can be either directly used or be adapted for the real dataset by the companies. In this paper, we study a new problem of synthesizing surrogate ER datasets using transformer models, with the goal that the ER model trained on the synthesized dataset can be used directly on the real dataset. We propose privacy preserving methods to synthesize ER datasets: we first learn the true similarity distributions of both matching and non-matching entity pairs from real dataset. We then devise algorithms that satisfy differential privacy and can synthesize fake but semantically meaningful entities, add matching and non-matching labels to these fake entity pairs, and ensure that the fake and real datasets have similar distributions. We also describe a method for entity rejection to avoid synthesizing bad fake entities that may destroy the original distributions. Extensive experiments show that ER matchers trained on real and synthetic ER datasets have very close performance on the same test sets - theirscores differ within 6% on 3 commonly used ER datasets, and their average precision, recall differences are less than 5%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Meta-Sim: Learning to Generate Synthetic DatasetsAmlan Kar, Aayush Prakash, Ming-Yu Liu, Eric Cameracci 等ICCV 2019 · 被引用 272 次
- ZeroER: Entity Resolution using Zero Labeled ExamplesRenzhi Wu, Sanya Chaba, Saurabh Sawlani, Xu Chu 等SIGMOD 2020 · 被引用 77 次
- Neural Data Server: A Large-Scale Search Engine for Transfer Learning DataXi Yan, David Acuna, Sanja FidlerCVPR 2020
- DatasetGAN: Efficient Labeled Data Factory With Minimal Human EffortYuxuan Zhang, Huan Ling, Jun Gao, Kangxue Yin 等CVPR 2021
相关 Paper
- Domain Adaptation for Deep Entity ResolutionJianhong Tu, Ju Fan, Nan Tang, Peng Wang 等SIGMOD 2022 · 被引用 46 次
- A Critical Re-evaluation of Record Linkage Benchmarks for Learning-Based Matching AlgorithmsGeorge Papadakis, Nishadi Kirielle, Peter Christen, Themis PalpanasICDE 2024 · 被引用 8 次
- CampER: An Effective Framework for Privacy-Aware Deep Entity ResolutionYuxiang Guo, Lu Chen, Zhengjie Zhou, Baihua Zheng 等KDD 2023 · 被引用 7 次
- Effective Explanations for Entity Resolution ModelsTommaso Teofili, Donatella Firmani, Nick Koudas, Vincenzo Martello 等ICDE 2022 · 被引用 18 次
- The Inadequacy of Similarity-Based Privacy Metrics: Privacy Attacks Against "Truly Anonymous" Synthetic DatasetsGeorgi Ganev, Emiliano De CristofaroS&P 2025
