Keytar: Practical Keystroke Timing Attacks and Input Reconstruction
Mufan Qiu, Lihsuan Chuang, Dohhyun Kim, Huaizhi Qu, Tianlong Chen, Andrew Kwong
Abstract
Keystroke timing attacks have long been recognized as a serious security concern. Researchers have conjectured that an attacker who learns the amount of time that elapses between keystrokes on a computer keyboard can reconstruct the keys pressed by a victim typist. Given the severe implications of a successful keystroke timing attack, numerous published side-channel works have utilized keystroke timing extraction as a case study to illustrate the impact of various types of side-channel attacks. However, despite an abundance of works demonstrating extraction of inter-keystroke timings, it remains to be proven that input recovery is actually possible.
This paper bridges this long-standing gap in the literature and performs a comprehensive study on the feasibility of reconstructing typed input from inter-keystroke timings. We model input reconstruction as a machine translation task and fine-tune open-source Large Language Models (LLMs) with a curriculum learning strategy, leveraging their ability to utilize contextual information and incorporate semantic understanding into the reconstruction process. With this approach, we reconstruct typed input with a high degree of fidelity. Using the best reconstruction among the Top-5 predictions and a normalized edit distance threshold of 0.1 as the criterion for successful reconstruction, we achieve a success rate of 34.9%.
We also demonstrate input reconstruction under practical, real-world circumstances, where additional noise is introduced to the inter-keystroke timing traces. We conduct end-to-end cache attacks, both from native environments and from the Chrome browser, and quantify how the additional noise inherent to cache attacks affects the input recovery process. To obtain a sufficiently large dataset for training and finetuning the LLM for noisy traces extracted via cache-attacks, we replayed over 1.5 million typing samples from real human typists while performing cache attacks. We release and open-source this dataset, along with our code and checkpoint for reconstructing input, so that future works on keystroketiming attacks can rigorously and empirically evaluate their effectiveness.
We provide both our tools for training LLMs on inter-keystroke timings and our tools for reconstructing input from inter-keystroke timings as open source code. This will enable side-channel researchers to train and evaluate our models against inter-keystroke timings traces obtained from various types of side-channels, which may exhibit different noise patterns. We envision that future studies demonstrating side-channels that extract inter-keystroke timings will use our tool to empirically demonstrate input reconstruction from their attacks, rather than merely speculating that the traces they obtain through their side-channels are sufficient. We also make public our dataset of inter-keystroke timings of 1.5 million sentences obtained through Prime+probe attacks so as to facilitate future research on algorithms for input reconstruction.
In this paper, we make the following contributions:
• We empirically demonstrate, for the first time, complete reconstruction of a victim typist's input from interkeystroke timings.
• We demonstrate the first reconstruction of user input from inter-keystroke timings obtained through cache-attacks, both from a native environment and from JavaScript code running within a browser in section 4.
• We build a dataset collection system that replays typing datasets while simultaneously obtaining side-channel traces. We utilize this system to generate a largescale dataset of inter-keystroke timings extracted through Prime+Probe attacks. We make this dataset public to facilitate subsequent research on algorithms for input reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on50
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Spectre Attacks: Exploiting Speculative ExecutionPaul Kocher, Jann Horn, Anders Fogh, Daniel Genkin et al.S&P 2019 · 2,435 citations
- Meltdown: Reading Kernel Memory from User SpaceMoritz Lipp, Michael Schwarz, Daniel Gruss, Thomas Prescher et al.USENIX Security 2018 · 1,456 citations
- Foreshadow: Extracting the Keys to the Intel SGX Kingdom with Transient Out-of-Order ExecutionJo Van Bulck, Marina Minkin, Ofir Weisse, Daniel Genkin et al.USENIX Security 2018 · 1,175 citations
- DRAMA: Exploiting DRAM Addressing for Cross-CPU AttacksPeter Pessl, Daniel Gruss, Clémentine Maurice, Michael Schwarz et al.USENIX Security 2016 · 500 citations
Related papers
- KeyDrown: Eliminating Software-Based Keystroke Timing Side-Channel AttacksMichael Schwarz, Moritz Lipp, Daniel Gruss, Samuel Weiser et al.NDSS 2018 · 68 citations
- Hidden No More: Attacking and Defending Private Third-Party LLM InferenceRahul Krishna Thomas, Louai Zahran, Erica Choi, Akilesh Potti et al.ICML 2025
- Port Contention for Fun and ProfitAlejandro Cabrera Aldaya, Billy Bob Brumley, Sohaib ul Hassan, Cesar Pereida García et al.S&P 2019 · 240 citations
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM InferenceZhifan Luo, Shuo Shao, Su Zhang, Lijing Zhou et al.NDSS 2026 · 32 citations
- I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM ServingGuanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang et al.NDSS 2025
