KHyperLogLog: Estimating Reidentifiability and Joinability of Large Data at Scale
Pern Hui Chia, Damien Desfontaines, Irippuge Milinda Perera, Daniel Simmons-Marengo, Chao Li, Wei-Yen Day, Qiushi Wang, Miguel Guevara
摘要
Understanding the privacy relevant characteristics of data sets, such as reidentifiability and joinability, is crucial for data governance, yet can be difficult for large data sets. While computing the data characteristics by brute force is straightforward, the scale of systems and data collected by large organizations demands an efficient approach. We present KHyperLogLog (KHLL), an algorithm based on approximate counting techniques that can estimate the reidentifiability and joinability risks of very large databases using linear runtime and minimal memory. KHLL enables one to measure reidentifiability of data quantitatively, rather than based on expert judgement or manual reviews. Meanwhile, joinability analysis using KHLL helps ensure the separation of pseudonymous and identified data sets. We describe how organizations can use KHLL to improve protection of user privacy. The efficiency of KHLL allows one to schedule periodic analyses that detect any deviations from the expected risks over time as a regression test for privacy. We validate the performance and accuracy of KHLL through experiments using proprietary and publicly available data sets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Discovering Related Data At ScaleSagar Bharadwaj, Praveen Gupta, Ranjita Bhagwan, Saikat GuhaVLDB 2021 · 被引用 23 次
- Discovering Similarity Inclusion DependenciesYouri Kaminsky, Eduardo H. M. Pena, Felix NaumannSIGMOD 2023 · 被引用 16 次
- Maximum Coverage in Turnstile Streams with Applications to Fingerprinting MeasuresAlina Ene, Alessandro Epasto, Vahab Mirrokni, Hoai-An Nguyen 等ICML 2025
相关 Paper
- Memory-Efficient Key/Foreign-Key Join Size Estimation via Multiplicity and Intersection SizeMagnus Müller, Daniel Flachs, Guido MoerkotteICDE 2021 · 被引用 4 次
- Nearly-Linear Time and Massively Parallel Algorithms for -anonymityKevin Aydin, Honghao Lin, David P. Woodruff, Peilin ZhongNeurIPS 2025
- Unmasking Vulnerabilities: Cardinality Sketches under Adaptive InputsSara Ahmadian, Edith CohenICML 2024 · 被引用 7 次
- UltraLogLog: A Practical and More Space-Efficient Alternative to HyperLogLog for Approximate Distinct CountingOtmar ErtlVLDB 2024 · 被引用 12 次
- CARBINE: Exploring Additional Properties of HyperLogLog for Secure and Robust Flow Cardinality EstimationDamu DingINFOCOM 2024 · 被引用 3 次
