Synthesizing Privacy Preserving Entity Resolution Datasets
Xuedi Qin, Chengliang Chai, Nan Tang, Jian Li, Yuyu Luo, Guoliang Li, Yaoyu Zhu
Abstract
Entity resolution (ER) is a core problem in data integration. Many companies have lots of datasets where ER needs to be conducted to integrate the data. On the one hand, it is nontrivial for non-ER experts within companies to design ER solutions. On the other hand, most companies are reluctant to release their real datasets for multiple reasons (e.g., privacy issues). A typical solution from the machine learning (ML) and the statistical community is to create surrogate (a.k.a. analogous) datasets based on the real dataset, release these surrogate datasets to the public to train ML models, such that these models trained on surrogate datasets can be either directly used or be adapted for the real dataset by the companies. In this paper, we study a new problem of synthesizing surrogate ER datasets using transformer models, with the goal that the ER model trained on the synthesized dataset can be used directly on the real dataset. We propose privacy preserving methods to synthesize ER datasets: we first learn the true similarity distributions of both matching and non-matching entity pairs from real dataset. We then devise algorithms that satisfy differential privacy and can synthesize fake but semantically meaningful entities, add matching and non-matching labels to these fake entity pairs, and ensure that the fake and real datasets have similar distributions. We also describe a method for entity rejection to avoid synthesizing bad fake entities that may destroy the original distributions. Extensive experiments show that ER matchers trained on real and synthetic ER datasets have very close performance on the same test sets - theirscores differ within 6% on 3 commonly used ER datasets, and their average precision, recall differences are less than 5%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89f79e5f-dc5a-4d3d-a7c6-35ab20b2994aCited by top-tier papers1
Ask how each one uses itBuilds on6
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Meta-Sim: Learning to Generate Synthetic DatasetsAmlan Kar, Aayush Prakash, Ming-Yu Liu, Eric Cameracci et al.ICCV 2019 · 272 citations
- ZeroER: Entity Resolution using Zero Labeled ExamplesRenzhi Wu, Sanya Chaba, Saurabh Sawlani, Xu Chu et al.SIGMOD 2020 · 77 citations
- Neural Data Server: A Large-Scale Search Engine for Transfer Learning DataXi Yan, David Acuna, Sanja FidlerCVPR 2020
- DatasetGAN: Efficient Labeled Data Factory With Minimal Human EffortYuxuan Zhang, Huan Ling, Jun Gao, Kangxue Yin et al.CVPR 2021
Related papers
- Domain Adaptation for Deep Entity ResolutionJianhong Tu, Ju Fan, Nan Tang, Peng Wang et al.SIGMOD 2022 · 46 citations
- A Critical Re-evaluation of Record Linkage Benchmarks for Learning-Based Matching AlgorithmsGeorge Papadakis, Nishadi Kirielle, Peter Christen, Themis PalpanasICDE 2024 · 8 citations
- CampER: An Effective Framework for Privacy-Aware Deep Entity ResolutionYuxiang Guo, Lu Chen, Zhengjie Zhou, Baihua Zheng et al.KDD 2023 · 7 citations
- Effective Explanations for Entity Resolution ModelsTommaso Teofili, Donatella Firmani, Nick Koudas, Vincenzo Martello et al.ICDE 2022 · 18 citations
- The Inadequacy of Similarity-Based Privacy Metrics: Privacy Attacks Against "Truly Anonymous" Synthetic DatasetsGeorgi Ganev, Emiliano De CristofaroS&P 2025
