Large-scale online deanonymization with LLMs
Simon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni, Nicholas Carlini, Florian Tramèr
摘要
We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer participants at high precision, given pseudonymous online profiles and conversations alone, matching what would take hours for a dedicated human investigator. We then design attacks for the closed-world setting. Given two databases of pseudonymous individuals, each containing unstructured text written by or about that individual, we implement a scalable attack pipeline that uses LLMs to: (1) extract identity-relevant features, (2) search for candidate matches via semantic embeddings, and (3) reason over top candidates to verify matches and reduce false positives. Compared to classical deanonymization work (e.g., on the Netflix prize) that required structured data, our approach works directly on raw user content across arbitrary platforms. We construct three datasets with known ground-truth data to evaluate our attacks. The first links Hacker News to LinkedIn profiles, using cross-platform references that appear in the profiles. Our second dataset matches users across Reddit movie discussion communities; and the third splits a single user's Reddit history in time to create two pseudonymous profiles to be matched. In each setting, LLM-based methods substantially outperform classical baselines, achieving up to 55% recall at 90% precision compared to near 0% for the best non-LLM method. Our results show that the cost of deanonymizing pseudonymous users online has fallen sharply, and threat models for online privacy should account for this shift.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 被引用 211 次
- Reducing Privacy Risks in Online Self-Disclosures with Language ModelsYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra 等ACL 2024 · 被引用 14 次
- Probabilistic Reasoning with LLMs for Privacy Risk EstimationJonathan Zheng, Alan Ritter, Sauvik Das, Wei (Coco) XuNeurIPS 2025 · 被引用 3 次
- Adversaries Can Misuse Combinations of Safe ModelsErik Jones, Anca D. Dragan, Jacob SteinhardtICML 2025
- Attacks on Deidentification's DefensesAloni CohenUSENIX Security 2022
相关 Paper
- De-Anonymization at Scale via Tournament-Style AttributionLirui Zhang, Huishuai ZhangACL 2026
- Language Models are Advanced AnonymizersRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2025
- Evaluating LLM-based Personal Information Extraction and CountermeasuresYupei Liu, Yuqi Jia, Jinyuan Jia, Neil Zhenqiang GongUSENIX Security 2025
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- Robust Utility-Preserving Text Anonymization Based on Large Language ModelsTianyu Yang, Xiaodan Zhu, Iryna GurevychACL 2025
