Anonymisation Models for Text Data: State of the art, Challenges and Future Directions
Pierre Lison, Ildikó Pilán, David Sánchez, Montserrat Batet, Lilja Øvrelid
Abstract
This position paper investigates the problem of automated text anonymisation, which is a prerequisite for secure sharing of documents containing sensitive information about individuals. We summarise the key concepts behind text anonymisation and provide a review of current approaches. Anonymisation methods have so far been developed in two fields with little mutual interaction, namely natural language processing and privacy-preserving data publishing. Based on a case study, we outline the benefits and limitations of these approaches and discuss a number of open challenges, such as (1) how to account for multiple types of semantic inferences, (2) how to strike a balance between disclosure risk and data utility and (3) how to evaluate the quality of the resulting anonymisation. We lay out a case for moving beyond sequence labelling models and incorporate explicit measures of disclosure risk into the text anonymisation process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0729fb24-b53a-4a13-8f19-d51e35b2631bCited by top-tier papers13
- "It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational AgentsZhiping Zhang, Michelle Jia, Hao-Ping (Hank) Lee, Bingsheng Yao et al.CHI 2024 · 92 citations
- Learning to Unlearn: Instance-Wise Unlearning for Pre-trained ClassifiersSungmin Cha, Sungjun Cho, Dasol Hwang, Honglak Lee et al.AAAI 2024 · 79 citations
- Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsJoel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha et al.ACL 2023 · 48 citations
- Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based ChatbotsJijie Zhou, Eryue Xu, Yaoyao Wu, Tianshi LiCHI 2025 · 15 citations
- Reducing Privacy Risks in Online Self-Disclosures with Language ModelsYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra et al.ACL 2024 · 14 citations
Related papers
- Robust Utility-Preserving Text Anonymization Based on Large Language ModelsTianyu Yang, Xiaodan Zhu, Iryna GurevychACL 2025
- Supporting Informed Self-Disclosure: Design Recommendations for Presenting AI-Estimates of Privacy Risks to UsersIsadora Krsek, Meryl Ye, Wei Xu, Alan Ritter et al.CHI 2026 · 1 citation
- How Private are Language Models in Abstractive Summarization?Anthony Hughes, Nikolaos Aletras, Ning MaEMNLP 2025
- ALSA: Context-Sensitive Prompt Privacy Preservation in Large Language ModelsHongru Ma, Wenpeng Lu, Yanjie Liang, Tianyi Wang et al.KDD 2025 · 1 citation
- Subject-level Inference for Realistic Text Anonymization EvaluationMyeong Seok Oh, Dong-Yun Kim, Hanseok Oh, Chaean Kang et al.ACL 2026
