Deduplicated Sampling On-Demand
Luca Zecchini, Vasilis Efthymiou, Felix Naumann, Giovanni Simonini
Abstract
Data practitioners often sample their datasets to produce representative subsets for their downstream tasks. When entities in a dataset can be partitioned into multiple groups, stratified sampling is commonly used to produce subsets that match a target group distribution, e.g., to select a balanced subset for training a machine learning model. However, real-world data frequently contains duplicates — multiple representations of the same real-world entity — that can bias sampling, necessitating deduplication.
We define deduplicated sampling as the task of producing a clean sample of a dirty dataset according to a target group distribution. The naïve approach to deduplicated sampling would first deduplicate the entire dataset upfront, then perform sampling ex post. However, that approach might be prohibitively expensive for large datasets and time/resource constraints. Deduplicated sampling ondemand with RadlER is a novel approach to produce a clean sample by focusing the cleaning effort only on entities required to appear in that sample. Our experimental evaluation, performed on multiple datasets from different domains, demonstrates that RadlER consistently outperforms baseline approaches, providing data scientists with an efficient solution to quickly produce a clean sample of a dirty dataset according to a target group distribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50c55816-7047-467f-b035-ca0555dd6bcbBuilds on16
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 325 citations
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani et al.VLDB 2021 · 109 citations
- Looking for Trouble: Analyzing Classifier Behavior via Pattern DivergenceEliana Pastor, Luca de Alfaro, Elena BaralisSIGMOD 2021 · 51 citations
- Tailoring Data Source Distributions for Fairness-aware Data IntegrationFatemeh Nargesian, Abolfazl Asudeh, H. V. JagadishVLDB 2021 · 51 citations
Related papers
- Entity Resolution On-DemandGiovanni Simonini, Luca Zecchini, Sonia Bergamaschi, Felix NaumannVLDB 2022 · 33 citations
- MDedup: Duplicate Detection with Matching DependenciesIoannis K. Koumarelas, Thorsten Papenbrock, Felix NaumannVLDB 2020 · 18 citations
- DUEL: Duplicate Elimination on Active Memory for Self-Supervised Class-Imbalanced LearningWon-Seok Choi, Hyundo Lee, Dong-Sig Han, Junseok Park et al.AAAI 2024 · 4 citations
- A Comprehensive Benchmark Framework for Active Learning Methods in Entity MatchingVenkata Vamsikrishna Meduri, Lucian Popa, Prithviraj Sen, Mohamed SarwatSIGMOD 2020 · 50 citations
- How do Categorical Duplicates Affect ML? A New Benchmark and Empirical AnalysesVraj Shah, Thomas J. Parashos, Arun KumarVLDB 2024 · 8 citations
