Watch and Crack: Password Inference from Smart-Glasses Video
Yoav Orenbach, Avishai Wool
摘要
Typing passwords on smartphones in public places exposes users to video-based side-channel attacks, in which an adversary records the typing session and reconstructs the entered password by analyzing finger movements. Prior video-based keystroke inference attacks have targeted free-form text on tablets, numeric PINs, pattern locks, and passwords under unrealistically simplified conditions.
In this paper we present the first general video-based keystroke inference attack pipeline that recovers rule-based alphanumeric passwords from smartphone QWERTY keyboards, covering all four keyboard layouts, and requiring no assumptions about the victim or their device. Our pipeline uses the built-in camera of the popular Meta Ray-Ban smart glasses to record a victim typing a password on a smartphone in a public setting. It tracks the device and typing fingertips at sub-pixel resolution, predicts keystrokes via a selfsupervised ensemble of neural networks, dynamically estimates the on-screen keyboard layout, and computes per-key probability distributions. These video-derived probabilities are then combined with prior password-typing distributions based on 2.19 billion leaked passwords to estimate the rank, and therefore the crack time, of the observed password.
We evaluate our pipeline via a user study, against passwords that adhere to a typical enterprise password policy, including both human-chosen and password-manager-chosen random passwords. Users typed naturally, some using two thumbs and some using the index finger. With our video side channel, human-chosen passwords of up to 16 characters can be cracked in hours to days, and password-manager-chosen random passwords of up to 13 characters can be cracked in under an hour on a modern GPU, reducing password entropy by up to 60 bits relative to the prior baseline. For human-chosen passwords, combining video evidence with prior statistics yields an additional statistically significant reduction of approximately 20 bits compared to either source alone. The attack succeeds at distances up to 1.8 m, across diverse attacker-victim postures (seated or standing) and viewing angles (from frontal to side-profile), and generalizes across three popular smartphones (Google Pixel 10, iPhone 16, and Samsung Galaxy A53).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath 等ICLR 2026 · 被引用 1,103 次
- Fast, Lean, and Accurate: Modeling Password Guessability Using Neural NetworksWilliam Melicher, Blase Ur, Sean M. Segreti, Saranga Komanduri 等USENIX Security 2016 · 被引用 331 次
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani 等ICLR 2021 · 被引用 294 次
- When CSI Meets Public WiFi: Inferring Your Mobile Phone Password via WiFi SignalsMengyuan Li, Yan Meng, Junyi Liu, Haojin Zhu 等CCS 2016 · 被引用 213 次
- Cracking Android Pattern Lock in Five AttemptsGuixin Ye, Zhanyong Tang, Dingyi Fang, Xiaojiang Chen 等NDSS 2017 · 被引用 123 次
相关 Paper
- Towards a General Video-based Keystroke Inference AttackZhuolin Yang, Yuxin Chen, Zain Sarwar, Hadleigh Schwartz 等USENIX Security 2023
- I Know Your Keyboard Input: A Robust Keystroke Eavesdropper Based-on Acoustic SignalsJia-Xuan Bai, Bin Liu, Luchuan SongACM MM 2021 · 被引用 24 次
- Eavesdropping user credentials via GPU side channels on smartphonesBoyuan Yang, Ruirong Chen, Kai Huang, Jun Yang 等ASPLOS 2022 · 被引用 12 次
- EyeTell: Video-Assisted Touchscreen Keystroke Inference from Eye MovementsYimin Chen, Tao Li, Rui Zhang, Yanchao Zhang 等S&P 2018 · 被引用 60 次
- A Keylogging Inference Attack on Air-Tapping Keyboards in Virtual EnvironmentsÜlkü Meteriz-Yildiran, Necip Fazil Yildiran, Amro Awad, David MohaisenIEEE VR 2022 · 被引用 40 次
