Discovering Significant Patterns under Sequential False Discovery Control
Sebastian Dalleiger, Jilles Vreeken
Abstract
We are interested in discovering those patterns from data with an empirical frequency that is significantly differently than expected. To avoid spurious results, yet achieve high statistical power, we propose to sequentially control for false discoveries during the search. To avoid redundancy, we propose to update our expectations whenever we discover a significant pattern. To efficiently consider the exponentially sized search space, we employ an easy-to-compute upper bound on significance, and propose an effective search strategy for sets of significant patterns. Through an extensive set of experiments on synthetic data, we show that our method, Spass, recovers the ground truth reliably, does so efficiently, and without redundancy. On real-world data we show it works well on both single and multiple classes, on low and high dimensional data, and through case studies that it discovers meaningful results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d65e3d6-c06d-4c45-976d-8b2e47aab753Cited by top-tier papers3
- Efficient Centrality Maximization with Rademacher AveragesLeonardo PellegrinaKDD 2023 · 9 citations
- Efficient Discovery of Significant Patterns with Few-Shot ResamplingLeonardo Pellegrina, Fabio VandinVLDB 2024 · 1 citation
- Few-Shot Resampling for Scalable Statistically-Sound Data MiningLeonardo Pellegrina, Fabio VandinKDD 2026
Builds on2
Related papers
- Statistically Significant Pattern Mining with Ordinal UtilityThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaKDD 2020 · 7 citations
- MCRapper: Monte-Carlo Rademacher Averages for Poset Families and Approximate Pattern MiningLeonardo Pellegrina, Cyrus Cousins, Fabio Vandin, Matteo RiondatoKDD 2020 · 7 citations
- Efficiently Enumerating Substrings with Statistically Significant Frequencies of Locally Optimal Occurrences in Gigantic StringAtsuyoshi Nakamura, Ichigaku Takigawa, Hiroshi MamitsukaAAAI 2020 · 2 citations
- Differentiable Pattern Set MiningJonas Fischer, Jilles VreekenKDD 2021 · 10 citations
- Discovering Approximate Functional Dependencies using Smoothed Mutual InformationFrédéric Pennerath, Panagiotis Mandros, Jilles VreekenKDD 2020 · 13 citations
