Lune

S&P2026顶会

Can Foundation LLMs Accurately Estimate Password Strength and Provide Appropriate Password Feedback?

Madison Pickering, Garrison Hinson-Hasty, Luca Dovichi, Helena Williams, Nathaniel Kim, Aybala Esmer, Blase Ur

2026年份

摘要

Password-strength models help users choose strong passwords. These models typically identify predictable passwords by training on millions of stolen passwords. We instead investigate if pretrained large language models (LLMs) can assess a password's strength based on semantic properties. We adapt LLMs in two distinct ways. First, mirroring LLMs' typical usage, we directly prompt four LLMs to estimate password strength. Second, mirroring prior probabilistic password models, we explicitly calculate the probability of passwords using lLMs fine-tuned on password data. We uncover technical challenges with doing so, developing strategies to reduce computational costs by a median of 25−97%25-97 \% depending on a password's length. We also perform an ablation study to gauge how key design dimensions (e.g., pretraining, data quantity) impact fine-tuned LLMs' accuracy. We find many fine-tuned models outperform prompted models despite being far smaller. Across both prompting and fine-tuning, LLMs modeled wordlike passwords well-including for out-of-distribution passwords (e.g., with recent cultural references)-yet struggled on more complex passwords. We also prompted LLMs to suggest how to improve specific passwords. The LLMs often gave poor or predictable feedback (e.g., making common substitutions, appending an exclamation point, or inserting “blue”).

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖