Measuring Re-identification Risk
CJ Carey, Travis Dick, Alessandro Epasto, Adel Javanmard, Josh Karlin, Shankar Kumar, Andres Muñoz Medina, Vahab Mirrokni, Gabriel Henrique Nunes, Sergei Vassilvitskii, Peilin Zhong
Abstract
Compact user representations (such as embeddings) form the backbone of personalization services. In this work, we present a new theoretical framework to measure re-identification risk in such user representations. Our framework, based on hypothesis testing, formally bounds the probability that an attacker may be able to obtain the identity of a user from their representation. As an application, we show how our framework is general enough to model important real-world applications such as the Chrome's Topics API for interest-based advertising. We complement our theoretical bounds by showing provably good attack algorithms for re-identification that we use to estimate the re-identification risk in the Topics API. We believe this work provides a rigorous and interpretable notion of re-identification risk and a framework to measure it that can be used to inform real-world applications.
- Also affiliated with Data Sciences and Operations, University of Southern California. † Also affiliated with Universidade Federal de Minas Gerais.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89b63601-1937-43c8-ab4c-c33f4cd30730Cited by top-tier papers5
- The Privacy-Utility Trade-off in the Topics APIMário S. Alvim, Natasha Fernandes, Annabelle McIver, Gabriel H. NunesCCS 2024 · 3 citations
- Auditing Privacy Mechanisms via Label Inference AttacksRóbert Busa-Fekete, Travis Dick, Claudio Gentile, Andrés Muñoz Medina et al.NeurIPS 2024 · 3 citations
- Understanding Disclosure Risk in Differential Privacy with Applications to Noise Calibration and AuditingPatricia Guerra-Balboa, Annika Sauer, Héber Hwang Arcolezi, Thorsten StrufeVLDB 2026
- Exploiting the Shared Storage APIAlexandra Nisenoff, Deian Stefan, Nicolas ChristinCCS 2025
- The Pervasive Blind Spot: Benchmarking VLM Inference Risks on Everyday Personal VideosShuning Zhang, Zhaoxin Li, Changxi Wen, Ying Ma et al.UbiComp 2026
Builds on2
Related papers
- Membership Inference Attacks Against Recommender SystemsMinxing Zhang, Zhaochun Ren, Zihan Wang, Pengjie Ren et al.CCS 2021 · 62 citations
- PassREfinder: Credential Stuffing Risk Prediction by Representing Password Reuse between Websites on a GraphJaehan Kim, Minkyoo Song, Minjae Seo, Youngjin Jin et al.S&P 2024 · 8 citations
- Targeted Deanonymization via the Cache Side Channel: Attacks and DefensesMojtaba Zaheri, Yossi Oren, Reza CurtmolaUSENIX Security 2022
- Free for All! Assessing User Data Exposure to Advertising Libraries on AndroidSoteris Demetriou, Whitney Merrill, Wei Yang, Aston Zhang et al.NDSS 2016 · 95 citations
- Unifying Re-Identification, Attribute Inference, and Data Reconstruction Risks in Differential PrivacyBogdan Kulynych, Juan Felipe Gómez, Georgios Kaissis, Jamie Hayes et al.NeurIPS 2025 · 15 citations
