How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses
Vraj Shah, Thomas J. Parashos, Arun Kumar
Abstract
The tedious grunt work involved in data preparation (prep) before ML reduces ML user productivity. It is also a roadblock to industrial-scale cloud AutoML workflows that build ML models for millions of datasets. One important data prep step for ML is cleaning duplicates in the Categorical columns, e.g., deduplicating CA with California in a State column. However, how such Categorical duplicates impact ML is ill-understood as there exist almost no in-depth scientific studies to assess their significance. In this work, we take the first step towards empirically characterizing the impact of Categorical duplicates on ML classification with a three-pronged approach. We first study how Categorical duplicates exhibit themselves by creating a labeled dataset of 1262 Categorical columns. We then curate a downstream benchmark suite of 16 real-world datasets to make observations on the effect of Categorical duplicates on five popular classifiers and five encoding mechanisms. We finally use simulation studies to validate our observations. We find that Logistic Regression and Similarity encoding are more robust to Categorical duplicates than two One-hot encoded high-capacity classifiers. We provide actionable takeaways that can potentially help AutoML developers to build better platforms and ML practitioners to reduce grunt work. While some of the presented insights have remained folklore for practitioners, our work presents the first systematic scientific study to analyze the impact of Categorical duplicates on ML and put this on an empirically rigorous footing. Our work presents novel data artifacts and benchmarks, as well as novel empirical analyses to spur more research on this topic.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 325 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Whither AutoML? Understanding the Role of Automation in Machine Learning WorkflowsDoris Xin, Eva Yiwei Wu, Doris Jung Lin Lee, Niloufar Salehi et al.CHI 2021 · 103 citations
- RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data PreparationNan Tang, Ju Fan, Fangyi Li, Jianhong Tu et al.VLDB 2021 · 92 citations
Related papers
- Towards Benchmarking Feature Type Inference for AutoML PlatformsVraj Shah, Jonathan Lacanlale, Premanand Kumar, Kevin Yang et al.SIGMOD 2021 · 16 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- A Large-Scale Study of Model Integration in ML-Enabled Software SystemsYorick Sens, Henriette Knopp, Sven Peldszus, Thorsten BergerICSE 2025 · 3 citations
- MDedup: Duplicate Detection with Matching DependenciesIoannis K. Koumarelas, Thorsten Papenbrock, Felix NaumannVLDB 2020 · 18 citations
- Using Random Effects to Account for High-Cardinality Categorical Features and Repeated Measures in Deep Neural NetworksGiora Simchoni, Saharon RossetNeurIPS 2021 · 27 citations
