Shielding PII to Prevent Re-identification and Preserve Utility
Shuhao Liu, Wenfei Fan, Yijia Xu
Abstract
This paper addresses the challenge of protecting Personally Identifiable Information (PII) in textual data, identifying and anonymizing PII to ensure privacy and regulatory compliance, while preserving data utility. We model this bi-criteria optimization problem as a two-player Stackelberg game, where an attacker seeks to link anonymized data back to individuals and a protector anonymizes the data to prevent re-identification. We show that the problem is intractable. Thus we develop SHIELD, an attack-aware PII protection system that iteratively engages the protector and attacker to prevent both PII breaches and over-scrubbing. SHIELD integrates logical reasoning with machine learning to identify PII, and supports pluggable attackers for robustness against re-identification. It achieves a constant-factor approximation for utility loss while mitigating risk. Using synthetic and real-world datasets, we empirically show that SHIELD offers better privacy-utility trade-off than prior PII protection systems, while remaining efficient and scalable.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2cabb632-8b01-476b-995c-76c33c1c98efRelated papers
- Robust Utility-Preserving Text Anonymization Based on Large Language ModelsTianyu Yang, Xiaodan Zhu, Iryna GurevychACL 2025
- PII-Bench: Evaluating Query-Aware Privacy Protection SystemsHao Shen, Zhouhong Gu, Haokai Hong, Weili Han et al.ACL 2026
- Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model UtilityMartin Kuo, Jingyang Zhang, Jianyi Zhang, Minxue Tang et al.ICLR 2025
- Mitigating Cumulative Privacy Risk in Continual Information Sharing: A Dynamic Stackelberg Game ApproachYuzi Yi, Weixuan Wang, Yehong Luo, Jinqiao Shi et al.WWW 2026
- Subject-level Inference for Realistic Text Anonymization EvaluationMyeong Seok Oh, Dong-Yun Kim, Hanseok Oh, Chaean Kang et al.ACL 2026
