Explainable Data Decompositions
Sebastian Dalleiger, Jilles Vreeken
摘要
Our goal is to discover the components of a dataset, characterize why we deem these components, explain how these components are different from each other, as well as identify what properties they share among each other. As is usual, we consider regions in the data to be components if they show significantly different distributions. What is not usual, however, is that we parameterize these distributions with patterns that are informative for one or more components. We do so because these patterns allow us to characterize what is going on in our data as well as explain our decomposition. We define the problem in terms of a regularized maximum likelihood, in which we use the Maximum Entropy principle to model each data component with a set of patterns. As the search space is large and unstructured, we propose the deterministic DISC algorithm to efficiently discover high-quality decompositions via an alternating optimization approach. Empirical evaluation on synthetic and real-world data shows that DISC efficiently discovers meaningful components and accurately characterises these in easily understandable terms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Differentiable Pattern Set MiningJonas Fischer, Jilles VreekenKDD 2021 · 被引用 10 次
- Discovering Significant Patterns under Sequential False Discovery ControlSebastian Dalleiger, Jilles VreekenKDD 2022 · 被引用 9 次
- Differentially Describing Groups of GraphsCorinna Coupette, Sebastian Dalleiger, Jilles VreekenAAAI 2022 · 被引用 8 次
相关 Paper
- Label-Descriptive Patterns and Their Application to Characterizing Classification ErrorsMichael A. Hedderich, Jonas Fischer, Dietrich Klakow, Jilles VreekenICML 2022 · 被引用 14 次
- Explainable Mixture Models through Differentiable Rule LearningMatthias Wilms, Sascha Xu, Jilles VreekenICLR 2026
- Learning disentangled representations via product manifold projectionMarco Fumero, Luca Cosmo, Simone Melzi, Emanuele RodolàICML 2021 · 被引用 29 次
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel 等NeurIPS 2021 · 被引用 421 次
- DisDiff: Unsupervised Disentanglement of Diffusion Probabilistic ModelsTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2023 · 被引用 41 次
