Private Data Imputation
Abdelkarim Kati, Florian Kerschbaum, Marina Blanton
Abstract
Data imputation is an important data preparation task where the data analyst replaces missing or erroneous values to increase the expected accuracy of downstream analyses. The accuracy improvement of data imputation extends to private data analyses across distributed databases. However, existing data imputation methods violate the privacy of the data rendering the privacy protection in the downstream analyses obsolete. We conclude that private data analysis requires private data imputation. In this paper, we present the first optimized protocols for private data imputation. We consider the case of horizontally and vertically split data sets. Our optimization aims to reduce most of the computation to private set intersection (or at least oblivious programmable pseudo-random function) protocols which can be very efficiently computed. We show that private data imputation has -- on average across all evaluated datasets -- an accuracy advantage of 20% in case of vertically split data and 5% in case of horizontally split data over imputing data locally. In case of the worst data split we observed that imputing using our method resulted in an increase of up to 32.7 times in the quality of imputation over the vertically split data and 3.4 times in case of horizontally split data. Our protocols are very efficient and run in 2.4 seconds in case of vertically split data and 8.4 seconds in case of horizontally split data for 100,000 records evaluated in the 10 Gbps network setting, performing one data imputation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3beb80c3-6454-4b11-82ce-29f07c95f8f3Builds on6
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta et al.NeurIPS 2021 · 573 citations
- PSI from PaXoS: Fast, Malicious Private Set IntersectionBenny Pinkas, Mike Rosulek, Ni Trieu, Avishay YanaiEUROCRYPT 2020 · 198 citations
- VOLE-PSI: Fast OPRF and Circuit-PSI from Vector-OLEPeter Rindal, Phillipp SchoppmannEUROCRYPT 2021 · 159 citations
- SANNS: Scaling Up Secure Approximate k-Nearest Neighbors SearchHao Chen, Ilaria Chillotti, Yihe Dong, Oxana Poburinnaya et al.USENIX Security 2020
- Faster Secure Comparisons with Offline Phase for Efficient Private Set IntersectionFlorian Kerschbaum, Erik-Oliver Blass, Rasoul Akhavan MahdaviNDSS 2023
Related papers
- Private Collaborative Data Cleaning via Non-Equi PSIErik-Oliver Blass, Florian KerschbaumS&P 2023
- Learnable Prompt as Pseudo-Imputation: Rethinking the Necessity of Traditional EHR Data Imputation in Downstream Clinical PredictionWeibin Liao, Yinghao Zhu, Zhongji Zhang, Yuhang Wang et al.KDD 2025 · 2 citations
- Secure Parallel Computation on Privately Partitioned Data and ApplicationsNuttapong Attrapadung, Hiraku Morita, Kazuma Ohara, Jacob C. N. Schuldt et al.CCS 2022 · 7 citations
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 5 citations
- Shuffle-based Private Set Union: Faster and More SecureYanxue Jia, Shifeng Sun, Hong-Sheng Zhou, Jiajun Du et al.USENIX Security 2022
