Efficient Discovery of Significant Patterns with Few-Shot Resampling
Leonardo Pellegrina, Fabio Vandin
摘要
Significant pattern mining is a fundamental task in mining transactional data, requiring to identify patterns significantly associated with the value of a given feature, the target. In several applications, such as biomedicine, basket market analysis, and social networks, the goal is to discover patterns whose association with the target is defined with respect to an underlying population, or process, of which the dataset represents only a collection of observations, or samples. A natural way to capture the association of a pattern with the target is to consider its statistical significance , assessing its deviation from the (null) hypothesis of independence between the pattern and the target. While several algorithms have been proposed to find statistically significant patterns, it remains a computationally demanding task, and for complex patterns such as subgroups, no efficient solution exists.
We present FSR, an efficient algorithm to identify statistically significant patterns with rigorous guarantees on the probability of false discoveries. FSR builds on a novel general framework for mining significant patterns that captures some of the most commonly considered patterns, including itemsets, sequential patterns, and subgroups. FSR uses a small number of resampled datasets, obtained by assigning i.i.d. labels to each transaction, to rigorously bound the supremum deviation of a quality statistic measuring the significance of patterns. FSR builds on novel tight bounds on the supremum deviation that require to mine a small number of resampled datasets, while providing a high effectiveness in discovering significant patterns. As a test case, we consider significant subgroup mining, and our evaluation on several real datasets shows that FSR is effective in discovering significant subgroups, while requiring a small number of resampled datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Efficient Temporal Pattern Mining in Big Time Series Using Mutual InformationVan Long Ho, Nguyen Ho, Torben Bach PedersenVLDB 2022 · 被引用 17 次
- What's in the Box? Exploring the Inner Life of Neural Networks with Robust RulesJonas Fischer, Anna Oláh, Jilles VreekenICML 2021 · 被引用 11 次
- Discovering Significant Patterns under Sequential False Discovery ControlSebastian Dalleiger, Jilles VreekenKDD 2022 · 被引用 9 次
- Fast and Scalable Mining of Time Series Motifs with Probabilistic GuaranteesMatteo Ceccarello, Johann GamperVLDB 2022 · 被引用 9 次
- MCRapper: Monte-Carlo Rademacher Averages for Poset Families and Approximate Pattern MiningLeonardo Pellegrina, Cyrus Cousins, Fabio Vandin, Matteo RiondatoKDD 2020 · 被引用 7 次
相关 Paper
- Statistically Significant Pattern Mining with Ordinal UtilityThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaKDD 2020 · 被引用 7 次
- FLEXIS: FLEXible Frequent Subgraph Mining using Maximal Independent SetsAkshit Sharma, Sam Reinehr, Dinesh Mehta, Bo WuKDD 2025
- HOPS: Probabilistic Subtree Mining for Small and Large GraphsPascal Welke, Florian Seiffarth, Michael Kamp, Stefan WrobelKDD 2020 · 被引用 5 次
- Non-overlapped Frequency based Episode Significance under Markov Null ModelsAvinash Achar, Santhosh B. Gandreti, Subbayya Sastry PidaparthyKDD 2026
- Finding Good Subtrees for Constraint Optimization Problems Using Frequent Pattern MiningHongbo Li, Jimmy Lee, He Mi, Minghao YinAAAI 2020 · 被引用 6 次
