Looking for Trouble: Analyzing Classifier Behavior via Pattern Divergence
Eliana Pastor, Luca de Alfaro, Elena Baralis
摘要
Machine learning models may perform differently on different data subgroups, which we represent as itemsets (i.e., conjunctions of simple predicates). The identification of these critical data subgroups plays an important role in many applications, for example model validation and testing, or evaluation of model fairness. Typically, domain expert help is required to identify relevant (or sensitive) subgroups.
We propose the notion of divergence over itemsets as a measure of different classification behavior on data subgroups, and the use of frequent pattern mining techniques for their identification. A quantification of the contribution of different attribute values to divergence, based on the mathematical foundations provided by Shapley values, allows us to identify both critical and peculiar behaviors of attributes. Extended experiments show the effectiveness of the approach in identifying critical subgroup behaviors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- A Unified Interactive Model Evaluation for Classification, Object Detection, and Instance Segmentation in Computer VisionChangjian Chen, Yukai Guo, Fengyuan Tian, Shilong Liu 等IEEE VIS 2023 · 被引用 28 次
- Fairness-Aware Range Queries for Selecting Unbiased DataSuraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam DasICDE 2022 · 被引用 19 次
- Understanding the Black Box: A Deep Empirical Dive into Shapley Value Approximations for Tabular DataSuchit Gupte, John PaparrizosSIGMOD 2025 · 被引用 19 次
- Detection of Groups with Biased Representation in RankingJinyang Li, Yuval Moskovitch, H. V. JagadishICDE 2023 · 被引用 10 次
- CohortNet: Empowering Cohort Discovery for Interpretable Healthcare AnalyticsQingpeng Cai, Kaiping Zheng, H. V. Jagadish, Beng Chin Ooi 等VLDB 2024 · 被引用 10 次
相关 Paper
- A Hierarchical Approach to Anomalous Subgroup DiscoveryEliana Pastor, Elena Baralis, Luca de AlfaroICDE 2023 · 被引用 9 次
- "Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsHaoran Zhang, Harvineet Singh, Marzyeh Ghassemi, Shalmali JoshiICML 2023 · 被引用 37 次
- Feature Importance Disparities for Data Bias InvestigationsPeter W. Chang, Leor Fishman, Seth NeelICML 2024 · 被引用 3 次
- Detecting Interpretable Subgroup DriftsFlavio Giobergia, Eliana Pastor, Luca de Alfaro, Elena BaralisKDD 2025
- WeightedSHAP: analyzing and improving Shapley based feature attributionsYongchan Kwon, James Y. ZouNeurIPS 2022 · 被引用 60 次
