Shielding PII to Prevent Re-identification and Preserve Utility
Shuhao Liu, Wenfei Fan, Yijia Xu
摘要
This paper addresses the challenge of protecting Personally Identifiable Information (PII) in textual data, identifying and anonymizing PII to ensure privacy and regulatory compliance, while preserving data utility. We model this bi-criteria optimization problem as a two-player Stackelberg game, where an attacker seeks to link anonymized data back to individuals and a protector anonymizes the data to prevent re-identification. We show that the problem is intractable. Thus we develop SHIELD, an attack-aware PII protection system that iteratively engages the protector and attacker to prevent both PII breaches and over-scrubbing. SHIELD integrates logical reasoning with machine learning to identify PII, and supports pluggable attackers for robustness against re-identification. It achieves a constant-factor approximation for utility loss while mitigating risk. Using synthetic and real-world datasets, we empirically show that SHIELD offers better privacy-utility trade-off than prior PII protection systems, while remaining efficient and scalable.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Robust Utility-Preserving Text Anonymization Based on Large Language ModelsTianyu Yang, Xiaodan Zhu, Iryna GurevychACL 2025
- PII-Bench: Evaluating Query-Aware Privacy Protection SystemsHao Shen, Zhouhong Gu, Haokai Hong, Weili Han 等ACL 2026
- Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model UtilityMartin Kuo, Jingyang Zhang, Jianyi Zhang, Minxue Tang 等ICLR 2025
- Mitigating Cumulative Privacy Risk in Continual Information Sharing: A Dynamic Stackelberg Game ApproachYuzi Yi, Weixuan Wang, Yehong Luo, Jinqiao Shi 等WWW 2026
- Subject-level Inference for Realistic Text Anonymization EvaluationMyeong Seok Oh, Dong-Yun Kim, Hanseok Oh, Chaean Kang 等ACL 2026
