De-Anonymizing Text by Fingerprinting Language Generation
Zhen Sun, Roei Schuster, Vitaly Shmatikov
摘要
Components of machine learning systems are not (yet) perceived as security hotspots. Secure coding practices, such as ensuring that no execution paths depend on confidential inputs, have not yet been adopted by ML developers. We initiate the study of code security of ML systems by investigating how nucleus sampling---a popular approach for generating text, used for applications such as auto-completion---unwittingly leaks texts typed by users. Our main result is that the series of nucleus sizes for many natural English word sequences is a unique fingerprint. We then show how an attacker can infer typed text by measuring these fingerprints via a suitable side channel (e.g., cache access times), explain how this attack could help de-anonymize anonymous texts, and discuss defenses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Deep Fingerprinting: Undermining Website Fingerprinting Defenses with Deep LearningPayap Sirinam, Mohsen Imani, Marc Juarez, Matthew WrightCCS 2018 · 被引用 632 次
- ARMageddon: Cache Attacks on Mobile DevicesMoritz Lipp, Daniel Gruss, Raphael Spreitzer, Clémentine Maurice 等USENIX Security 2016 · 被引用 451 次
- Translation Leak-aside Buffer: Defeating Cache Side-channel Protections with TLB AttacksBen Gras, Kaveh Razavi, Herbert Bos, Cristiano GiuffridaUSENIX Security 2018 · 被引用 357 次
- Rendered Insecure: GPU Side Channel Attacks are PracticalHoda Naghibijouybari, Ajaya Neupane, Zhiyun Qian, Nael B. Abu-GhazalehCCS 2018 · 被引用 214 次
相关 Paper
- There's always a bigger fish: a clarifying analysis of a machine-learning-assisted side-channel attackJack Cook, Jules Drean, Jonathan Behrens, Mengjia YanISCA 2022 · 被引用 35 次
- When Coding Style Survives Compilation: De-anonymizing Programmers from Executable BinariesAylin Caliskan, Fabian Yamaguchi, Edwin Dauber, Richard E. Harang 等NDSS 2018 · 被引用 125 次
- Traces of Memorisation in Large Language Models for CodeAli Al-Kaswan, Maliheh Izadi, Arie van DeursenICSE 2024 · 被引用 23 次
- Black-Box Adversarial Attacks on LLM-Based Code CompletionSlobodan Jenko, Niels Mündler, Jingxuan He, Mark Vero 等ICML 2025
- Unveiling your keystrokes: A Cache-based Side-channel Attack on Graphics LibrariesDaimeng Wang, Ajaya Neupane, Zhiyun Qian, Nael B. Abu-Ghazaleh 等NDSS 2019 · 被引用 47 次
