A hierarchical decomposition for explaining ML performance discrepancies
Harvineet Singh, Fan Xia, Adarsh Subbaswamy, Alexej Gossmann, Jean Feng
Abstract
Machine learning (ML) algorithms can often differ in performance across domains. Understanding their performance differs is crucial for determining what types of interventions (e.g., algorithmic or operational) are most effective at closing the performance gaps. Existing methods focus on of the total performance gap into the impact of a shift in the distribution of features versus the impact of a shift in the conditional distribution of the outcome ; however, such coarse explanations offer only a few options for how one can close the performance gap. that quantify the importance of each variable to each term in the aggregate decomposition can provide a much deeper understanding and suggest much more targeted interventions. However, existing methods assume knowledge of the full causal graph or make strong parametric assumptions. We introduce a nonparametric hierarchical framework that provides both aggregate and detailed decompositions for explaining why the performance of an ML algorithm differs across domains, without requiring causal knowledge. We derive debiased, computationally-efficient estimators, and statistical inference procedures for asymptotically valid confidence intervals.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82c9cc4b-3216-4f98-a635-cfbd8c353b80Cited by top-tier papers3
- "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance DriftHarvineet Singh, Fan Xia, Alexej Gossmann, Andrew Chuang et al.ICML 2025
- Explaining Concept Shift with Interpretable Feature AttributionRuiqi Lyu, Alistair Turcan, Bryan WilderICML 2026
- Going Beyond Static: Understanding Shifts with Time-Series AttributionJiashuo Liu, Nabeel Seedat, Peng Cui, Mihaela van der SchaarICLR 2025
Builds on5
- Efficient nonparametric statistical inference on population feature importance using Shapley valuesBrian D. Williamson, Jean FengICML 2020 · 86 citations
- Feature Shift Detection: Localizing Which Features Have Shifted via Conditional Distribution TestsSean Kulinski, Saurabh Bagchi, David I. InouyeNeurIPS 2020 · 39 citations
- Towards Explaining Distribution ShiftsSean Kulinski, David I. InouyeICML 2023 · 38 citations
- "Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsHaoran Zhang, Harvineet Singh, Marzyeh Ghassemi, Shalmali JoshiICML 2023 · 37 citations
- Distilling Model Failures as Directions in Latent SpaceSaachi Jain, Hannah Lawrence, Ankur Moitra, Aleksander MadryICLR 2023 · 11 citations
Related papers
- Explaining Algorithmic Fairness Through Fairness-Aware Causal Path DecompositionWeishen Pan, Sen Cui, Jiang Bian, Changshui Zhang et al.KDD 2021 · 29 citations
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras et al.NeurIPS 2025 · 9 citations
- Debiasing Concept-based Explanations with Causal AnalysisMohammad Taha Bahadori, David HeckermanICLR 2021 · 8 citations
- Returning The Favour: When Regression Benefits From Probabilistic Causal KnowledgeShahine Bouabid, Jake Fawkes, Dino SejdinovicICML 2023
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
