Subject-level Inference for Realistic Text Anonymization Evaluation
Myeong Seok Oh, Dong-Yun Kim, Hanseok Oh, Chaean Kang, Joeun Kang, Xiaonan Wang, Hyunjung Park, Young Cheol Jung, Hansaem Kim
Abstract
Current text anonymization evaluation relies on span-based metrics that fail to capture what an adversary could actually infer, and assumes a single data subject, ignoring multi-subject scenarios. To address these limitations, we present SPIA (Subject-level PII Inference Assessment), the first benchmark that shifts the unit of evaluation from text spans to individuals, comprising 675 documents across legal and online domains with novel subject-level protection metrics. Extensive experiments show that even when over 90% of PII spans are masked, subject-level inference protection drops as low as 33%, leaving the majority of personal information recoverable through contextual inference. Furthermore, target-subject-focused anonymization leaves non-target subjects substantially more exposed than the target subject. We show that subject-level inferencebased evaluation is essential for ensuring safe text anonymization in real-world settings. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68257d5a-8e8f-4efa-87aa-a133d85ec09eBuilds on5
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- MaSS: Multi-attribute Selective Suppression for Utility-preserving Data Transformation from an Information-theoretic PerspectiveYizhuo Chen, Chun-Fu Chen, Hsiang Hsu, Shaohan Hu et al.ICML 2024 · 3 citations
- Anonymisation Models for Text Data: State of the art, Challenges and Future DirectionsPierre Lison, Ildikó Pilán, David Sánchez, Montserrat Batet et al.ACL 2021
- PII-Bench: Evaluating Query-Aware Privacy Protection SystemsHao Shen, Zhouhong Gu, Haokai Hong, Weili Han et al.ACL 2026
- Analyzing Leakage of Personally Identifiable Information in Language ModelsNils Lukas, Ahmed Salem, Robert Sim, Shruti Tople et al.S&P 2023
Related papers
- Language Models are Advanced AnonymizersRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2025
- RedacBench: Can AI Erase Your Secrets?Hyunjun Jeon, Kyuyoung Kim, Jinwoo ShinICLR 2026 · 2 citations
- Concept-Aware Privacy Mechanisms for Defending Embedding Inversion AttacksYu-Che Tsai, Hsiang Hsiao, Kuan-Yu Chen, Shou-De LinICLR 2026 · 2 citations
- AnonPSI: An Anonymity Assessment Framework for PSIBo Jiang, Jian Du, Qiang YanNDSS 2024
- Shielding PII to Prevent Re-identification and Preserve UtilityShuhao Liu, Wenfei Fan, Yijia XuSIGMOD 2026
