Subject-level Inference for Realistic Text Anonymization Evaluation
Myeong Seok Oh, Dong-Yun Kim, Hanseok Oh, Chaean Kang, Joeun Kang, Xiaonan Wang, Hyunjung Park, Young Cheol Jung, Hansaem Kim
摘要
Current text anonymization evaluation relies on span-based metrics that fail to capture what an adversary could actually infer, and assumes a single data subject, ignoring multi-subject scenarios. To address these limitations, we present SPIA (Subject-level PII Inference Assessment), the first benchmark that shifts the unit of evaluation from text spans to individuals, comprising 675 documents across legal and online domains with novel subject-level protection metrics. Extensive experiments show that even when over 90% of PII spans are masked, subject-level inference protection drops as low as 33%, leaving the majority of personal information recoverable through contextual inference. Furthermore, target-subject-focused anonymization leaves non-target subjects substantially more exposed than the target subject. We show that subject-level inferencebased evaluation is essential for ensuring safe text anonymization in real-world settings. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- MaSS: Multi-attribute Selective Suppression for Utility-preserving Data Transformation from an Information-theoretic PerspectiveYizhuo Chen, Chun-Fu Chen, Hsiang Hsu, Shaohan Hu 等ICML 2024 · 被引用 3 次
- Anonymisation Models for Text Data: State of the art, Challenges and Future DirectionsPierre Lison, Ildikó Pilán, David Sánchez, Montserrat Batet 等ACL 2021
- PII-Bench: Evaluating Query-Aware Privacy Protection SystemsHao Shen, Zhouhong Gu, Haokai Hong, Weili Han 等ACL 2026
- Analyzing Leakage of Personally Identifiable Information in Language ModelsNils Lukas, Ahmed Salem, Robert Sim, Shruti Tople 等S&P 2023
相关 Paper
- Language Models are Advanced AnonymizersRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2025
- RedacBench: Can AI Erase Your Secrets?Hyunjun Jeon, Kyuyoung Kim, Jinwoo ShinICLR 2026 · 被引用 2 次
- Concept-Aware Privacy Mechanisms for Defending Embedding Inversion AttacksYu-Che Tsai, Hsiang Hsiao, Kuan-Yu Chen, Shou-De LinICLR 2026 · 被引用 2 次
- AnonPSI: An Anonymity Assessment Framework for PSIBo Jiang, Jian Du, Qiang YanNDSS 2024
- Shielding PII to Prevent Re-identification and Preserve UtilityShuhao Liu, Wenfei Fan, Yijia XuSIGMOD 2026
