Differentiable Pattern Set Mining
Jonas Fischer, Jilles Vreeken
Abstract
Pattern set mining has been successful in discovering small sets of highly informative and useful patterns from data. To find good models, existing methods heuristically explore the twice-exponential search space over all possible pattern sets in a combinatorial way, by which they are limited to data over at most hundreds of features, as well as likely to get stuck in local minima. Here, we propose a gradient based optimization approach that allows us to efficiently discover high-quality pattern sets from data of millions of rows and hundreds of thousands of features.
In particular, we propose a novel type of neural autoencoder called BinaPs, using binary activations and binarizing weights in each forward pass, which are directly interpretable as conjunctive patterns. For training, optimizing a data-sparsity aware reconstruction loss, continuous versions of the weights are learned in small, noisy steps. This formulation provides a link between the discrete search space and continuous optimization, thus allowing for a gradient based strategy to discover sets of high-quality and noise-robust patterns. Through extensive experiments on both synthetic and real world data, we show that BinaPs discovers high quality and noise robust patterns, and unique among all competitors, easily scales to data of supermarket transactions or biological variant calls.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f53e1f8c-0d2c-4b77-90b2-ee2f1f3362eeCited by top-tier papers4
- Efficiently Factorizing Boolean Matrices using Proximal Gradient DescentSebastian Dalleiger, Jilles VreekenNeurIPS 2022 · 8 citations
- Learning Exceptional Subgroups by End-to-End Maximizing KL-DivergenceSascha Xu, Nils Philipp Walter, Janis Kalofolias, Jilles VreekenICML 2024 · 8 citations
- Language-Model Based Informed Partition of Databases to Speed Up Pattern MiningCarlos Bobed Lisbona, Jordi Bernad, Pierre MaillotSIGMOD 2024
- Stochastic Submodular Data ForgettingRamón Rico, Arno Siebes, Yannis VelegrakisSIGMOD 2026
Builds on2
Related papers
- Finding Interpretable Class-Specific Patterns through Efficient Neural SearchNils Philipp Walter, Jonas Fischer, Jilles VreekenAAAI 2024 · 8 citations
- Discovering Significant Patterns under Sequential False Discovery ControlSebastian Dalleiger, Jilles VreekenKDD 2022 · 9 citations
- Differentiably Discovering Sets of RulesLuis N. J. Paulus, Jonas Fischer, Jilles VreekenKDD 2026
- Towards Reliable Neural SpecificationsChuqin Geng, Nham Le, Xiaojie Xu, Zhaoyue Wang et al.ICML 2023 · 14 citations
- Neural-based classification rule learning for sequential dataMarine Collery, Philippe Bonnard, François Fages, Remy KustersICLR 2023
