USENIX Security2026Top-tier venue
Large-scale online deanonymization with LLMs
Simon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni, Nicholas Carlini, Florian Tramèr
Abstract
We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer participants at high precision, given pseudonymous online profiles and conversations alone, matching what would take hours for a dedicated human investigator. We then design attacks for the closed-world setting. Given two databases of pseudonymous individuals, each containing unstructured text written by or about that individual, we implement a scalable attack pipeline that uses LLMs to: (1) extract identity-relevant features, (2) search for candidate matches via semantic embeddings, and (3) reason over top candidates to verify matches and reduce false positives. Compared to classical deanonymization work (e.g., on the Netflix prize) that required structured data, our approach works directly on raw user content across arbitrary platforms. We construct three datasets with known ground-truth data to evaluate our attacks. The first links Hacker News to LinkedIn profiles, using cross-platform references that appear in the profiles. Our second dataset matches users across Reddit movie discussion communities; and the third splits a single user's Reddit history in time to create two pseudonymous profiles to be matched. In each setting, LLM-based methods substantially outperform classical baselines, achieving up to 55% recall at 90% precision compared to near 0% for the best non-LLM method. Our results show that the cost of deanonymizing pseudonymous users online has fallen sharply, and threat models for online privacy should account for this shift.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c3b9cff-08d4-438b-9005-c0250e574b77Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 211 citations
- Reducing Privacy Risks in Online Self-Disclosures with Language ModelsYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra et al.ACL 2024 · 14 citations
- Probabilistic Reasoning with LLMs for Privacy Risk EstimationJonathan Zheng, Alan Ritter, Sauvik Das, Wei (Coco) XuNeurIPS 2025 · 3 citations
- Adversaries Can Misuse Combinations of Safe ModelsErik Jones, Anca D. Dragan, Jacob SteinhardtICML 2025
- Attacks on Deidentification's DefensesAloni CohenUSENIX Security 2022
Related papers
- De-Anonymization at Scale via Tournament-Style AttributionLirui Zhang, Huishuai ZhangACL 2026
- Language Models are Advanced AnonymizersRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2025
- Evaluating LLM-based Personal Information Extraction and CountermeasuresYupei Liu, Yuqi Jia, Jinyuan Jia, Neil Zhenqiang GongUSENIX Security 2025
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 291 citations
- Robust Utility-Preserving Text Anonymization Based on Large Language ModelsTianyu Yang, Xiaodan Zhu, Iryna GurevychACL 2025
