Looking for Trouble: Analyzing Classifier Behavior via Pattern Divergence
Eliana Pastor, Luca de Alfaro, Elena Baralis
Abstract
Machine learning models may perform differently on different data subgroups, which we represent as itemsets (i.e., conjunctions of simple predicates). The identification of these critical data subgroups plays an important role in many applications, for example model validation and testing, or evaluation of model fairness. Typically, domain expert help is required to identify relevant (or sensitive) subgroups.
We propose the notion of divergence over itemsets as a measure of different classification behavior on data subgroups, and the use of frequent pattern mining techniques for their identification. A quantification of the contribution of different attribute values to divergence, based on the mathematical foundations provided by Shapley values, allows us to identify both critical and peculiar behaviors of attributes. Extended experiments show the effectiveness of the approach in identifying critical subgroup behaviors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac9cee10-40dd-48a3-b151-9e293beac376Cited by top-tier papers15
- A Unified Interactive Model Evaluation for Classification, Object Detection, and Instance Segmentation in Computer VisionChangjian Chen, Yukai Guo, Fengyuan Tian, Shilong Liu et al.IEEE VIS 2023 · 28 citations
- Fairness-Aware Range Queries for Selecting Unbiased DataSuraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam DasICDE 2022 · 19 citations
- Understanding the Black Box: A Deep Empirical Dive into Shapley Value Approximations for Tabular DataSuchit Gupte, John PaparrizosSIGMOD 2025 · 19 citations
- Detection of Groups with Biased Representation in RankingJinyang Li, Yuval Moskovitch, H. V. JagadishICDE 2023 · 10 citations
- CohortNet: Empowering Cohort Discovery for Interpretable Healthcare AnalyticsQingpeng Cai, Kaiping Zheng, H. V. Jagadish, Beng Chin Ooi et al.VLDB 2024 · 10 citations
Related papers
- A Hierarchical Approach to Anomalous Subgroup DiscoveryEliana Pastor, Elena Baralis, Luca de AlfaroICDE 2023 · 9 citations
- "Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsHaoran Zhang, Harvineet Singh, Marzyeh Ghassemi, Shalmali JoshiICML 2023 · 37 citations
- Feature Importance Disparities for Data Bias InvestigationsPeter W. Chang, Leor Fishman, Seth NeelICML 2024 · 3 citations
- Detecting Interpretable Subgroup DriftsFlavio Giobergia, Eliana Pastor, Luca de Alfaro, Elena BaralisKDD 2025
- WeightedSHAP: analyzing and improving Shapley based feature attributionsYongchan Kwon, James Y. ZouNeurIPS 2022 · 60 citations
