De-Anonymizing Text by Fingerprinting Language Generation
Zhen Sun, Roei Schuster, Vitaly Shmatikov
Abstract
Components of machine learning systems are not (yet) perceived as security hotspots. Secure coding practices, such as ensuring that no execution paths depend on confidential inputs, have not yet been adopted by ML developers. We initiate the study of code security of ML systems by investigating how nucleus sampling---a popular approach for generating text, used for applications such as auto-completion---unwittingly leaks texts typed by users. Our main result is that the series of nucleus sizes for many natural English word sequences is a unique fingerprint. We then show how an attacker can infer typed text by measuring these fingerprints via a suitable side channel (e.g., cache access times), explain how this attack could help de-anonymize anonymous texts, and discuss defenses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cb79d29-ccc4-4661-932a-dfbc47e7de90Cited by top-tier papers1
Ask how each one uses itBuilds on11
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Deep Fingerprinting: Undermining Website Fingerprinting Defenses with Deep LearningPayap Sirinam, Mohsen Imani, Marc Juarez, Matthew WrightCCS 2018 · 632 citations
- ARMageddon: Cache Attacks on Mobile DevicesMoritz Lipp, Daniel Gruss, Raphael Spreitzer, Clémentine Maurice et al.USENIX Security 2016 · 451 citations
- Translation Leak-aside Buffer: Defeating Cache Side-channel Protections with TLB AttacksBen Gras, Kaveh Razavi, Herbert Bos, Cristiano GiuffridaUSENIX Security 2018 · 357 citations
- Rendered Insecure: GPU Side Channel Attacks are PracticalHoda Naghibijouybari, Ajaya Neupane, Zhiyun Qian, Nael B. Abu-GhazalehCCS 2018 · 214 citations
Related papers
- There's always a bigger fish: a clarifying analysis of a machine-learning-assisted side-channel attackJack Cook, Jules Drean, Jonathan Behrens, Mengjia YanISCA 2022 · 35 citations
- When Coding Style Survives Compilation: De-anonymizing Programmers from Executable BinariesAylin Caliskan, Fabian Yamaguchi, Edwin Dauber, Richard E. Harang et al.NDSS 2018 · 125 citations
- Traces of Memorisation in Large Language Models for CodeAli Al-Kaswan, Maliheh Izadi, Arie van DeursenICSE 2024 · 23 citations
- Black-Box Adversarial Attacks on LLM-Based Code CompletionSlobodan Jenko, Niels Mündler, Jingxuan He, Mark Vero et al.ICML 2025
- Unveiling your keystrokes: A Cache-based Side-channel Attack on Graphics LibrariesDaimeng Wang, Ajaya Neupane, Zhiyun Qian, Nael B. Abu-Ghazaleh et al.NDSS 2019 · 47 citations
