Anonymisation Models for Text Data: State of the art, Challenges and Future Directions
Pierre Lison, Ildikó Pilán, David Sánchez, Montserrat Batet, Lilja Øvrelid
摘要
This position paper investigates the problem of automated text anonymisation, which is a prerequisite for secure sharing of documents containing sensitive information about individuals. We summarise the key concepts behind text anonymisation and provide a review of current approaches. Anonymisation methods have so far been developed in two fields with little mutual interaction, namely natural language processing and privacy-preserving data publishing. Based on a case study, we outline the benefits and limitations of these approaches and discuss a number of open challenges, such as (1) how to account for multiple types of semantic inferences, (2) how to strike a balance between disclosure risk and data utility and (3) how to evaluate the quality of the resulting anonymisation. We lay out a case for moving beyond sequence labelling models and incorporate explicit measures of disclosure risk into the text anonymisation process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- "It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational AgentsZhiping Zhang, Michelle Jia, Hao-Ping (Hank) Lee, Bingsheng Yao 等CHI 2024 · 被引用 92 次
- Learning to Unlearn: Instance-Wise Unlearning for Pre-trained ClassifiersSungmin Cha, Sungjun Cho, Dasol Hwang, Honglak Lee 等AAAI 2024 · 被引用 79 次
- Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsJoel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha 等ACL 2023 · 被引用 48 次
- Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based ChatbotsJijie Zhou, Eryue Xu, Yaoyao Wu, Tianshi LiCHI 2025 · 被引用 15 次
- Reducing Privacy Risks in Online Self-Disclosures with Language ModelsYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra 等ACL 2024 · 被引用 14 次
相关 Paper
- Robust Utility-Preserving Text Anonymization Based on Large Language ModelsTianyu Yang, Xiaodan Zhu, Iryna GurevychACL 2025
- Supporting Informed Self-Disclosure: Design Recommendations for Presenting AI-Estimates of Privacy Risks to UsersIsadora Krsek, Meryl Ye, Wei Xu, Alan Ritter 等CHI 2026 · 被引用 1 次
- How Private are Language Models in Abstractive Summarization?Anthony Hughes, Nikolaos Aletras, Ning MaEMNLP 2025
- ALSA: Context-Sensitive Prompt Privacy Preservation in Large Language ModelsHongru Ma, Wenpeng Lu, Yanjie Liang, Tianyi Wang 等KDD 2025 · 被引用 1 次
- Subject-level Inference for Realistic Text Anonymization EvaluationMyeong Seok Oh, Dong-Yun Kim, Hanseok Oh, Chaean Kang 等ACL 2026
