Data-Constrained Synthesis of Training Data for De-Identification
Thomas Vakili, Aron Henriksson, Hercules Dalianis
Abstract
Many sensitive domains -- such as the clinical domain -- lack widely available datasets due to privacy risks. The increasing generative capabilities of large language models (LLMs) have made synthetic datasets a viable path forward. In this study, we domain-adapt LLMs to the clinical domain and generate synthetic clinical texts that are machine-annotated with tags for personally identifiable information using capable encoder-based NER models. The synthetic corpora are then used to train synthetic NER models. The results show that training NER models using synthetic corpora incurs only a small drop in predictive performance. The limits of this process are investigated in a systematic ablation study -- using both Swedish and Spanish data. Our analysis shows that smaller datasets can be sufficient for domain-adapting LLMs for data synthesis. Instead, the effectiveness of this process is almost entirely contingent on the performance of the machine-annotating NER models trained using the original data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aeed09c6-7de1-42cc-9dad-035cfa74bd0dBuilds on5
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar et al.ACL 2023 · 24 citations
- S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation ExtractionBenfeng Xu, Quan Wang, Yajuan Lyu, Dai Dai et al.ACL 2023 · 13 citations
Related papers
- How Private are Language Models in Abstractive Summarization?Anthony Hughes, Nikolaos Aletras, Ning MaEMNLP 2025
- Expert-guided Clinical Text Augmentation via Query-Based Model CollaborationDongkyu Cho, Miao Zhang, Gregory Lyng, Rumi ChunaraICML 2026
- Enhancing Small Medical Learners with Privacy-preserving Contextual PromptingXinlu Zhang, Shiyang Li, Xianjun Yang, Chenxin Tian et al.ICLR 2024 · 13 citations
- Synthetic Text Generation for Training Large Language Models via Gradient MatchingDang Nguyen, Zeman Li, MohammadHossein Bateni, Vahab Mirrokni et al.ICML 2025
- Can Personal Health Information Be Secured in LLM? Privacy Attack and Defense in the Medical DomainYujin Kang, Eunsun Kim, Yoon-Sik ChoCCS 2025
