Targeted Password Guessing Using k-Nearest Neighbors
Zhen Li, Ding Wang
Abstract
As the number of users’ password accounts are constantly increasing, users are more and more inclined to reuse passwords. Recently, considerable efforts have been made to construct targeted password guessing models to characterize users’ password reuse behaviors. However, existing studies mainly focus on characterizing slight modifications by training only on similar password pairs (e.g., Shark0301 → shark03). This leads to overfitting and causes existing models to overlook users’ large modification behaviors (e.g., Shark0301 → Bear03). To fill this gap, this paper introduces a new non-parametric method named k-nearest-neighbors targeted password guessing (KNNTPG). KNN-TPG builds a datastore that retains the context vector of all source passwords along with prefixes of the targeted passwords. During the generation of a new password, KNNTPG retrieves k nearest neighbor vectors from the datastore to ensure that the generated passwords align better with realistic password distributions. By creatively combining KNN-TPG with our proposed Transformer-based password model, we propose a new targeted password guessing model, namely KNNGuess. At each step of generating a new password, KNNGuess predicts and utilizes three distinct distributions, aiming to comprehensively model users’ password reuse behaviors. We demonstrate the effectiveness of our KNNGuess model and the KNN-TPG method through extensive experiments, which include 12 large-scale real-world password datasets, containing 4.8 billion passwords. More specifically, when the victim’s password at site A is compromised (namely pwA), within 100 guesses, the cracking success rate of KNNGuess for guessing her password at site B (namely pwB , and pwB ̸=pwA) is 25.40% (for common users) and 10.26% (for security-savvy users), which is 8.52%119.0% (avg. 55.33%) higher than its foremost counterparts. When comparing with state-of-the-art password models (i.e., Pass2Edit and PointerGuess), this value is 8.52%-27.66% (avg. 18.09%) higher. Our results highlight that the threat of password tweaking attacks is higher than users expected.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8d2191b-1b61-41a9-bc3e-2f63c65aa77bBuilds on19
- Targeted Online Password Guessing: An Underestimated ThreatDing Wang, Zijian Zhang, Ping Wang, Jeff Yan et al.CCS 2016 · 385 citations
- Fast, Lean, and Accurate: Modeling Password Guessability Using Neural NetworksWilliam Melicher, Blase Ur, Sean M. Segreti, Saranga Komanduri et al.USENIX Security 2016 · 331 citations
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2021 · 323 citations
- How I Learned to be Secure: a Census-Representative Survey of Security Advice Sources and BehaviorElissa M. Redmiles, Sean Kross, Michelle L. MazurekCCS 2016 · 192 citations
- Let's Go in for a Closer Look: Observing Passwords in Their Natural HabitatSarah Pearman, Jeremy Thomas, Pardis Emami Naeini, Hana Habib et al.CCS 2017 · 168 citations
Related papers
- PointerGuess: Targeted Password Guessing Model Using Pointer MechanismKedong Xiu, Ding WangUSENIX Security 2024 · 12 citations
- Pass2Edit: A Multi-Step Generative Model for Guessing Edited PasswordsDing Wang, Yunkai Zou, Yuan-an Xiao, Siqi Ma et al.USENIX Security 2023
- Improving Real-world Password Guessing Attacks via Bi-directional TransformersMing Xu, Jitao Yu, Xinyi Zhang, Chuanwang Wang et al.USENIX Security 2023
- Beyond Credential Stuffing: Password Similarity Models Using Neural NetworksBijeeta Pal, Tal Daniel, Rahul Chatterjee, Thomas RistenpartS&P 2019 · 100 citations
- RankGuess: Password Guessing Using Adversarial RankingTao Yang, Ding WangS&P 2025
