Parallel Rule Discovery from Large Datasets by Sampling
Wenfei Fan, Ziyan Han, Yaoshu Wang, Min Xie
Abstract
Rule discovery from large datasets is often prohibitively costly. The problem becomes more staggering when the rules are collectively defined across multiple tables. To scale with large datasets, this paper proposes a multi-round sampling strategy for rule discovery. We consider entity enhancing rules (REEs) for collective entity resolution and conflict resolution, which may carry constant patterns and machine learning predicates. We sample large datasets with accuracy bounds a and B such that at least a% of rules discovered from samples are guaranteed to hold on the entire dataset (i.e., precision), and at least B% of rules on the entire dataset can be mined from the samples (i.e., recall). We also quantify the connection between support and confidence of the rules on samples and their counterparts on the entire dataset. To scale with the number of tuple variables in collective rules, we adopt deep Q-learning to select semantically relevant predicates. To improve the recall, we develop a tableau method to recover constant patterns from the dataset. We parallelize the algorithm such that it guarantees to reduce runtime when more processors are used. Using real-life and synthetic data, we empirically verify that the method speeds up REE discovery by 12.2 times with sample ratio 10% and recall 82%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79745c26-4a19-48ec-b347-e7a7584894c9Cited by top-tier papers9
- Discovering Top-k Rules using Subjective and Objective CriteriaWenfei Fan, Ziyan Han, Yaoshu Wang, Min XieSIGMOD 2023 · 10 citations
- Splitting Tuples of Mismatched EntitiesWenfei Fan, Ziyan Han, Weilong Ren, Ding Wang et al.SIGMOD 2024 · 5 citations
- BClean: A Bayesian Data Cleaning SystemJianbin Qin, Sifan Huang, Yaoshu Wang, Jing Zhu et al.ICDE 2024 · 4 citations
- Discovering Top-k Relevant and Diversified RulesWenfei Fan, Ziyan Han, Min Xie, Guangyi ZhangSIGMOD 2025 · 2 citations
- Capturing More Associations by Referencing External GraphsWenfei Fan, Muyang Liu, Shuhao Liu, Chao TianVLDB 2024 · 2 citations
Builds on9
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Discovery of Approximate (and Exact) Denial ConstraintsEduardo H. M. Pena, Eduardo C. de Almeida, Felix NaumannVLDB 2020 · 79 citations
- Approximate Denial ConstraintsEster Livshits, Alireza Heidari, Ihab F. Ilyas, Benny KimelfeldVLDB 2020 · 60 citations
- Improving the Efficiency and Effectiveness for BERT-based Entity ResolutionBing Li, Yukai Miao, Yaoshu Wang, Yifang Sun et al.AAAI 2021 · 45 citations
- A Statistical Perspective on Discovering Functional Dependencies in Noisy DataYunjia Zhang, Zhihan Guo, Theodoros RekatsinasSIGMOD 2020 · 45 citations
Related papers
- Deep and Collective Entity Resolution in ParallelTing Deng, Wenfei Fan, Ping Lu, Xiaomeng Luo et al.ICDE 2022 · 6 citations
- Incremental Rule Discovery in Response to Parameter UpdatesHaoxian Chen, Wenfei Fan, Jiaye ZhengSIGMOD 2025 · 2 citations
- Making It Tractable to Catch Duplicates and Conflicts in GraphsWenfei Fan, Wenzhi Fu, Ruochun Jin, Muyang Liu et al.SIGMOD 2023 · 10 citations
- Parallel Discrepancy Detection and Incremental DetectionWenfei Fan, Chao Tian, Yanghao Wang, Qiang YinVLDB 2021 · 29 citations
- Discovering Association Rules from Big GraphsWenfei Fan, Wenzhi Fu, Ruochun Jin, Ping Lu et al.VLDB 2022 · 29 citations
