Differentially Private Password Frequency Lists
Jeremiah Blocki, Anupam Datta, Joseph Bonneau
摘要
Given a dataset of user-chosen passwords, the frequency list reveals the frequency of each unique password. We present a novel mechanism for releasing perturbed password frequency lists with rigorous security, efficiency, and distortion guarantees. Specifically, our mechanism is based on a novel algorithm for sampling that enables an efficient implementation of the exponential mechanism for differential privacy (naive sampling is exponential time). It provides the security guarantee that an adversary will not be able to use this perturbed frequency list to learn anything of significance about any individual user's password even if the adversary already possesses a wealth of background knowledge about the users in the dataset. We prove that our mechanism introduces minimal distortion, thus ensuring that the released frequency list is close to the actual list. Further, we empirically demonstrate, using the now-canonical password dataset leaked from RockYou, that the mechanism works well in practice: as the differential privacy parameter varies from 8 to 0.002 (smaller implies higher security), the normalized distortion coefficient (representing the distance between the released and actual password frequency list divided by the number of users N ) varies from 8.8 × 10 -7 to 1.9 × 10 -3 . Given this appealing combination of security and distortion guarantees, our mechanism enables organizations to publish perturbed password frequency lists. This can facilitate new research comparing password security between populations and evaluating password improvement approaches. To this end, we have collaborated with Yahoo! to use our differentially private mechanism to publicly release a corpus of 50 password frequency lists representing approximately 70 million Yahoo! users. This dataset is now the largest password frequency corpus available. Using our perturbed dataset we are able to closely replicate the original published analysis of this data. Permission to freely reproduce all or part of this paper for noncommercial purposes is granted provided that copies bear this notice and the full citation on the first page. Reproduction for commercial purposes is strictly prohibited without the prior written consent of the Internet Society, the first-named author (for reproduction of an entire paper only), and the author's employer if the paper was prepared within the scope of employment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Deep Models Under the GAN: Information Leakage from Collaborative Deep LearningBriland Hitaj, Giuseppe Ateniese, Fernando Pérez-CruzCCS 2017 · 被引用 1,581 次
- On the Economics of Offline Password CrackingJeremiah Blocki, Benjamin Harsha, Samson ZhouS&P 2018 · 被引用 79 次
- Secure Multi-party Computation of Differentially Private Heavy HittersJonas Böhler, Florian KerschbaumCCS 2021 · 被引用 34 次
- A Joint Exponential Mechanism For Differentially Private Top-kJennifer Gillenwater, Matthew Joseph, Andres Muñoz Medina, Mónica Ribero DiazICML 2022 · 被引用 20 次
- Optimizing Fitness-For-Use of Differentially Private Linear QueriesYingtai Xiao, Zeyu Ding, Yuxin Wang, Danfeng Zhang 等VLDB 2021 · 被引用 19 次
相关 Paper
- Towards a Rigorous Statistical Analysis of Empirical Password DatasetsJeremiah Blocki, Peiyuan LiuS&P 2023
- How to Attack and Generate HoneywordsDing Wang, Yunkai Zou, Qiying Dong, Yuanming Song 等S&P 2022 · 被引用 45 次
- Frequency Estimation under Local Differential PrivacyGraham Cormode, Samuel Maddock, Carsten MapleVLDB 2021 · 被引用 70 次
- Differentially-private software frequency profiling under linear constraintsHailong Zhang, Yu Hao, Sufian Latif, Raef Bassily 等OOPSLA 2020 · 被引用 1 次
- A Security Analysis of HoneywordsDing Wang, Haibo Cheng, Ping Wang, Jeff Yan 等NDSS 2018 · 被引用 1,102 次
