De-Identification of Sensitive Personal Data in Datasets Derived from IIT-CDIP
Stefan Larson, Nicole Lima, Santiago Diaz, Amogh Manoj Joshi, Siddharth Betala, Jamiu T. Suleiman, Yash Mathur, Kaushal Prajapati, Ramla Alakraa, Junjie Shen, Temi Okotore, Kevin Leach
Abstract
The IIT-CDIP document collection is the source of several widely used and publicly accessible document understanding datasets. In this paper, manual inspection of 5 datasets derived from IIT-CDIP uncovers the presence of thousands of instances of sensitive personal data, including US Social Security Numbers (SSNs), birth places and dates, and home addresses of individuals. The presence of such sensitive personal data in commonly-used and publicly available datasets is startling and has ethical and potentially legal implications; we believe such sensitive data ought to be removed from the internet. Thus, in this paper, we develop a modular data de-identification pipeline that replaces sensitive data with synthetic, but realistic, data. Via experiments, we demonstrate that this de-identification method preserves the utility of the de-identified documents so that they can continue be used in various document understanding applications. We will release redacted versions of these datasets publicly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
Related papers
- CodexLeaks: Privacy Leaks from Code Generation Language Models in GitHub CopilotLiang Niu, Muhammad Shujaat Mirza, Zayd Maradni, Christina PöpperUSENIX Security 2023
- On the Privacy Risks Caused by Images in Academic PapersSze Yiu Chau, Jianyu Lu, Kin Man Leung, Chi Fung Kwan et al.CCS 2026
- Inferring Users' Demographics and Sensitive Interests Using the Topics APIAthicha Srivirote, Muhammad Abu Bakar Aziz, Jeffrey L. Gleason, Desheng Hu et al.WWW 2026
- From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM AgentsMyeongseob Ko, Jihyun Jeong, Sumiran Thakur, Gyuhak Kim et al.ICML 2026 · 3 citations
- RedacBench: Can AI Erase Your Secrets?Hyunjun Jeon, Kyuyoung Kim, Jinwoo ShinICLR 2026 · 2 citations
