De-Identification of Sensitive Personal Data in Datasets Derived from IIT-CDIP
Stefan Larson, Nicole Lima, Santiago Diaz, Amogh Manoj Joshi, Siddharth Betala, Jamiu T. Suleiman, Yash Mathur, Kaushal Prajapati, Ramla Alakraa, Junjie Shen, Temi Okotore, Kevin Leach
摘要
The IIT-CDIP document collection is the source of several widely used and publicly accessible document understanding datasets. In this paper, manual inspection of 5 datasets derived from IIT-CDIP uncovers the presence of thousands of instances of sensitive personal data, including US Social Security Numbers (SSNs), birth places and dates, and home addresses of individuals. The presence of such sensitive personal data in commonly-used and publicly available datasets is startling and has ethical and potentially legal implications; we believe such sensitive data ought to be removed from the internet. Thus, in this paper, we develop a modular data de-identification pipeline that replaces sensitive data with synthetic, but realistic, data. Via experiments, we demonstrate that this de-identification method preserves the utility of the de-identified documents so that they can continue be used in various document understanding applications. We will release redacted versions of these datasets publicly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie 等ICCV 2021 · 被引用 392 次
相关 Paper
- CodexLeaks: Privacy Leaks from Code Generation Language Models in GitHub CopilotLiang Niu, Muhammad Shujaat Mirza, Zayd Maradni, Christina PöpperUSENIX Security 2023
- On the Privacy Risks Caused by Images in Academic PapersSze Yiu Chau, Jianyu Lu, Kin Man Leung, Chi Fung Kwan 等CCS 2026
- Inferring Users' Demographics and Sensitive Interests Using the Topics APIAthicha Srivirote, Muhammad Abu Bakar Aziz, Jeffrey L. Gleason, Desheng Hu 等WWW 2026
- From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM AgentsMyeongseob Ko, Jihyun Jeong, Sumiran Thakur, Gyuhak Kim 等ICML 2026 · 被引用 3 次
- RedacBench: Can AI Erase Your Secrets?Hyunjun Jeon, Kyuyoung Kim, Jinwoo ShinICLR 2026 · 被引用 2 次
