Deduplicated Sampling On-Demand
Luca Zecchini, Vasilis Efthymiou, Felix Naumann, Giovanni Simonini
摘要
Data practitioners often sample their datasets to produce representative subsets for their downstream tasks. When entities in a dataset can be partitioned into multiple groups, stratified sampling is commonly used to produce subsets that match a target group distribution, e.g., to select a balanced subset for training a machine learning model. However, real-world data frequently contains duplicates — multiple representations of the same real-world entity — that can bias sampling, necessitating deduplication.
We define deduplicated sampling as the task of producing a clean sample of a dirty dataset according to a target group distribution. The naïve approach to deduplicated sampling would first deduplicate the entire dataset upfront, then perform sampling ex post. However, that approach might be prohibitively expensive for large datasets and time/resource constraints. Deduplicated sampling ondemand with RadlER is a novel approach to produce a clean sample by focusing the cleaning effort only on entities required to appear in that sample. Our experimental evaluation, performed on multiple datasets from different domains, demonstrates that RadlER consistently outperforms baseline approaches, providing data scientists with an efficient solution to quickly produce a clean sample of a dirty dataset according to a target group distribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani 等VLDB 2021 · 被引用 109 次
- Looking for Trouble: Analyzing Classifier Behavior via Pattern DivergenceEliana Pastor, Luca de Alfaro, Elena BaralisSIGMOD 2021 · 被引用 51 次
- Tailoring Data Source Distributions for Fairness-aware Data IntegrationFatemeh Nargesian, Abolfazl Asudeh, H. V. JagadishVLDB 2021 · 被引用 51 次
相关 Paper
- Entity Resolution On-DemandGiovanni Simonini, Luca Zecchini, Sonia Bergamaschi, Felix NaumannVLDB 2022 · 被引用 33 次
- MDedup: Duplicate Detection with Matching DependenciesIoannis K. Koumarelas, Thorsten Papenbrock, Felix NaumannVLDB 2020 · 被引用 18 次
- DUEL: Duplicate Elimination on Active Memory for Self-Supervised Class-Imbalanced LearningWon-Seok Choi, Hyundo Lee, Dong-Sig Han, Junseok Park 等AAAI 2024 · 被引用 4 次
- A Comprehensive Benchmark Framework for Active Learning Methods in Entity MatchingVenkata Vamsikrishna Meduri, Lucian Popa, Prithviraj Sen, Mohamed SarwatSIGMOD 2020 · 被引用 50 次
- How do Categorical Duplicates Affect ML? A New Benchmark and Empirical AnalysesVraj Shah, Thomas J. Parashos, Arun KumarVLDB 2024 · 被引用 8 次
