Lune

S&P2026Top-tier venue

Can Foundation LLMs Accurately Estimate Password Strength and Provide Appropriate Password Feedback?

Madison Pickering, Garrison Hinson-Hasty, Luca Dovichi, Helena Williams, Nathaniel Kim, Aybala Esmer, Blase Ur

2026Year

Abstract

Password-strength models help users choose strong passwords. These models typically identify predictable passwords by training on millions of stolen passwords. We instead investigate if pretrained large language models (LLMs) can assess a password's strength based on semantic properties. We adapt LLMs in two distinct ways. First, mirroring LLMs' typical usage, we directly prompt four LLMs to estimate password strength. Second, mirroring prior probabilistic password models, we explicitly calculate the probability of passwords using lLMs fine-tuned on password data. We uncover technical challenges with doing so, developing strategies to reduce computational costs by a median of 25−97%25-97 \% depending on a password's length. We also perform an ablation study to gauge how key design dimensions (e.g., pretraining, data quantity) impact fine-tuned LLMs' accuracy. We find many fine-tuned models outperform prompted models despite being far smaller. Across both prompting and fine-tuning, LLMs modeled wordlike passwords well-including for out-of-distribution passwords (e.g., with recent cultural references)-yet struggled on more complex passwords. We also prompted LLMs to suggest how to improve specific passwords. The LLMs often gave poor or predictable feedback (e.g., making common substitutions, appending an exclamation point, or inserting “blue”).

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 697f028d-9267-41a2-8049-f8635b354f2b

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines