Attacks on Deidentification's Defenses
Aloni Cohen
摘要
Quasi-identifier-based deidentification techniques (QI-deidentification) are widely used in practice, including -anonymity, -diversity, and -closeness. We present three new attacks on QI-deidentification: two theoretical attacks and one practical attack on a real dataset. In contrast to prior work, our theoretical attacks work even if every attribute is a quasi-identifier. Hence, they apply to -anonymity, -diversity, -closeness, and most other QI-deidentification techniques. First, we introduce a new class of privacy attacks called downcoding attacks, and prove that every QI-deidentification scheme is vulnerable to downcoding attacks if it is minimal and hierarchical. Second, we convert the downcoding attacks into powerful predicate singling-out (PSO) attacks, which were recently proposed as a way to demonstrate that a privacy mechanism fails to legally anonymize under Europe's General Data Protection Regulation. Third, we use LinkedIn.com to reidentify 3 students in a -anonymized dataset published by EdX (and show thousands are potentially vulnerable), undermining EdX's claimed compliance with the Family Educational Rights and Privacy Act. The significance of this work is both scientific and political. Our theoretical attacks demonstrate that QI-deidentification may offer no protection even if every attribute is treated as a quasi-identifier. Our practical attack demonstrates that even deidentification experts acting in accordance with strict privacy regulations fail to prevent real-world reidentification. Together, they rebut a foundational tenet of QI-deidentification and challenge the actual arguments made to justify the continued use of -anonymity and other QI-deidentification techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- SoK: Privacy-Preserving Data SynthesisYuzheng Hu, Fan Wu, Qinbin Li, Yunhui Long 等S&P 2024 · 被引用 61 次
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 被引用 37 次
- On the Risks of Collecting Multidimensional Data Under Local Differential PrivacyHéber Hwang Arcolezi, Sébastien Gambs, Jean-François Couchot, Catuscia PalamidessiVLDB 2023 · 被引用 22 次
- Large-scale online deanonymization with LLMsSimon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni 等USENIX Security 2026 · 被引用 20 次
- SoK: Technical Implementation and Human Impact of Internet Privacy RegulationsEleanor Birrell, Jay Rodolitz, Angel Ding, Jenna Lee 等S&P 2024 · 被引用 11 次
相关 Paper
- Privacy-preserving datasets of eye-tracking samples with applications in XRBrendan David-John, Kevin R. B. Butler, Eakta JainIEEE VR 2023 · 被引用 35 次
- Targeted Deanonymization via the Cache Side Channel: Attacks and DefensesMojtaba Zaheri, Yossi Oren, Reza CurtmolaUSENIX Security 2022
- Re-identification Attack to Privacy-Preserving Data Analysis with Noisy Sample-MeanDu Su, Hieu Tri Huynh, Ziao Chen, Yi Lu 等KDD 2020 · 被引用 12 次
- Differential Privacy and Swapping: Examining De-Identification's Impact on Minority Representation and Privacy Preservation in the U.S. CensusMiranda Christ, Sarah Radway, Steven M. BellovinS&P 2022 · 被引用 23 次
- Exposing Privacy Risks in Anonymizing Clinical Data: Combinatorial Refinement Attacks on k-Anonymity Without Auxiliary InformationSomiya Chhillar, Mary K. Righi, Rebecca E. Sutter, Evgenios M. KornaropoulosCCS 2025
