Parallel Rule Discovery from Large Datasets by Sampling
Wenfei Fan, Ziyan Han, Yaoshu Wang, Min Xie
摘要
Rule discovery from large datasets is often prohibitively costly. The problem becomes more staggering when the rules are collectively defined across multiple tables. To scale with large datasets, this paper proposes a multi-round sampling strategy for rule discovery. We consider entity enhancing rules (REEs) for collective entity resolution and conflict resolution, which may carry constant patterns and machine learning predicates. We sample large datasets with accuracy bounds a and B such that at least a% of rules discovered from samples are guaranteed to hold on the entire dataset (i.e., precision), and at least B% of rules on the entire dataset can be mined from the samples (i.e., recall). We also quantify the connection between support and confidence of the rules on samples and their counterparts on the entire dataset. To scale with the number of tuple variables in collective rules, we adopt deep Q-learning to select semantically relevant predicates. To improve the recall, we develop a tableau method to recover constant patterns from the dataset. We parallelize the algorithm such that it guarantees to reduce runtime when more processors are used. Using real-life and synthetic data, we empirically verify that the method speeds up REE discovery by 12.2 times with sample ratio 10% and recall 82%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Discovering Top-k Rules using Subjective and Objective CriteriaWenfei Fan, Ziyan Han, Yaoshu Wang, Min XieSIGMOD 2023 · 被引用 10 次
- Splitting Tuples of Mismatched EntitiesWenfei Fan, Ziyan Han, Weilong Ren, Ding Wang 等SIGMOD 2024 · 被引用 5 次
- BClean: A Bayesian Data Cleaning SystemJianbin Qin, Sifan Huang, Yaoshu Wang, Jing Zhu 等ICDE 2024 · 被引用 4 次
- Discovering Top-k Relevant and Diversified RulesWenfei Fan, Ziyan Han, Min Xie, Guangyi ZhangSIGMOD 2025 · 被引用 2 次
- Capturing More Associations by Referencing External GraphsWenfei Fan, Muyang Liu, Shuhao Liu, Chao TianVLDB 2024 · 被引用 2 次
它引用的顶会 Paper9
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Discovery of Approximate (and Exact) Denial ConstraintsEduardo H. M. Pena, Eduardo C. de Almeida, Felix NaumannVLDB 2020 · 被引用 79 次
- Approximate Denial ConstraintsEster Livshits, Alireza Heidari, Ihab F. Ilyas, Benny KimelfeldVLDB 2020 · 被引用 60 次
- Improving the Efficiency and Effectiveness for BERT-based Entity ResolutionBing Li, Yukai Miao, Yaoshu Wang, Yifang Sun 等AAAI 2021 · 被引用 45 次
- A Statistical Perspective on Discovering Functional Dependencies in Noisy DataYunjia Zhang, Zhihan Guo, Theodoros RekatsinasSIGMOD 2020 · 被引用 45 次
相关 Paper
- Deep and Collective Entity Resolution in ParallelTing Deng, Wenfei Fan, Ping Lu, Xiaomeng Luo 等ICDE 2022 · 被引用 6 次
- Incremental Rule Discovery in Response to Parameter UpdatesHaoxian Chen, Wenfei Fan, Jiaye ZhengSIGMOD 2025 · 被引用 2 次
- Making It Tractable to Catch Duplicates and Conflicts in GraphsWenfei Fan, Wenzhi Fu, Ruochun Jin, Muyang Liu 等SIGMOD 2023 · 被引用 10 次
- Parallel Discrepancy Detection and Incremental DetectionWenfei Fan, Chao Tian, Yanghao Wang, Qiang YinVLDB 2021 · 被引用 29 次
- Discovering Association Rules from Big GraphsWenfei Fan, Wenzhi Fu, Ruochun Jin, Ping Lu 等VLDB 2022 · 被引用 29 次
