A Hierarchical Approach to Anomalous Subgroup Discovery
Eliana Pastor, Elena Baralis, Luca de Alfaro
Abstract
Understanding peculiar and anomalous behavior of machine learning models for specific data subgroups is a fundamental building block of model performance and fairness evaluation. The analysis of these data subgroups can provide useful insights into model inner working and highlight its potentially discriminatory behavior. Current approaches to subgroup exploration ignore the presence of hierarchies in the data, and can only be applied to discretized attributes. The discretization process required for continuous attributes may significantly affect the identification of relevant subgroups.We propose a hierarchical subgroup exploration technique to identify anomalous subgroup behavior at multiple granularity levels, along with a technique for the hierarchical discretization of data attributes. The hierarchical discretization produces, for each continuous attribute, a hierarchy of intervals. The subsequent hierarchical exploration can exploit data hierarchies, selecting for each attribute the optimal granularity to identify subgroups that are both anomalous, and with enough elements to be statistically and practically significant. Compared to non- hierarchical approaches, we show that our hierarchical approach is more powerful in identifying anomalous subgroups and more stable with respect to discretization and exploration parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bf03112-eee2-4ceb-8112-34d506455004Cited by top-tier papers3
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 24 citations
- Fair and Actionable Causal Prescription RulesetBenton Li, Nativ Levy, Brit Youngmann, Sainyam Galhotra et al.SIGMOD 2025 · 3 citations
- Detecting Interpretable Subgroup DriftsFlavio Giobergia, Eliana Pastor, Luca de Alfaro, Elena BaralisKDD 2025
Builds on4
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Looking for Trouble: Analyzing Classifier Behavior via Pattern DivergenceEliana Pastor, Luca de Alfaro, Elena BaralisSIGMOD 2021 · 51 citations
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 45 citations
- Identifying Insufficient Data Coverage for Ordinal Continuous-Valued AttributesAbolfazl Asudeh, Nima Shahbazi, Zhongjun Jin, H. V. JagadishSIGMOD 2021 · 30 citations
Related papers
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras et al.NeurIPS 2025 · 9 citations
- Subgroups Matter for Robust Bias MitigationAnissa Alloula, Charles Jones, Ben Glocker, Bartlomiej W. PapiezICML 2025
- Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for SupportMichael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan et al.CSCW 2022 · 149 citations
- Bounding and Approximating Intersectional Fairness through Marginal FairnessMathieu Molina, Patrick LoiseauNeurIPS 2022 · 16 citations
- Multi-group Learning for Hierarchical GroupsSamuel Deng, Daniel HsuICML 2024 · 7 citations
