"Why did the Model Fail?": Attributing Model Performance Changes to Distribution Shifts
Haoran Zhang, Harvineet Singh, Marzyeh Ghassemi, Shalmali Joshi
Abstract
Machine learning models frequently experience performance drops under distribution shifts. The underlying cause of such shifts may be multiple simultaneous factors such as changes in data quality, differences in specific covariate distributions, or changes in the relationship between label and features. When a model does fail during deployment, attributing performance change to these factors is critical for the model developer to identify the root cause and take mitigating actions. In this work, we introduce the problem of attributing performance differences between environments to distribution shifts in the underlying data generating mechanisms. We formulate the problem as a cooperative game where the players are distributions. We define the value of a set of distributions to be the change in model performance when only this set of distributions has changed between environments, and derive an importance weighting method for computing the value of an arbitrary set of distributions. The contribution of each distribution to the total performance change is then quantified as its Shapley value. We demonstrate the correctness and utility of our method on synthetic, semi-synthetic, and real-world case studies, showing its effectiveness in attributing performance changes to a wide range of distribution shifts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3b40a75-5ffd-4af2-8f78-1cc011681b9eCited by top-tier papers10
- Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World EnvironmentsPaulius Rauba, Nabeel Seedat, Krzysztof Kacprzyk, Mihaela van der SchaarNeurIPS 2024 · 15 citations
- A hierarchical decomposition for explaining ML performance discrepanciesHarvineet Singh, Fan Xia, Adarsh Subbaswamy, Alexej Gossmann et al.NeurIPS 2024 · 9 citations
- EigenScore: OOD Detection using Posterior Covariance in Diffusion ModelsShirin Shoushtari, Yi Wang, Xiao Shi, M. Salman Asif et al.ICLR 2026 · 5 citations
- ICYM2I: The illusion of multimodal informativeness under missingnessYoung Sang Choi, Vincent Jeanselme, Pierre A. Elias, Shalmali JoshiICLR 2026 · 2 citations
- Path-specific effects for pulse-oximetry guided decisions in critical careKevin Zhang, Yonghan Jung, Divyat Mahajan, Karthikeyan Shanmugam et al.NeurIPS 2025 · 2 citations
Builds on11
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 799 citations
- A Fine-Grained Analysis on Distribution ShiftOlivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre-Alvise Rebuffi et al.ICLR 2022 · 258 citations
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 149 citations
- Feature Shift Detection: Localizing Which Features Have Shifted via Conditional Distribution TestsSean Kulinski, Saurabh Bagchi, David I. InouyeNeurIPS 2020 · 39 citations
Related papers
- Explaining Probabilistic Models with Distributional ValuesLuca Franceschi, Michele Donini, Cédric Archambeau, Matthias W. SeegerICML 2024 · 4 citations
- Multiply-Robust Causal Change AttributionVictor Quintas-Martinez, Mohammad Taha Bahadori, Eduardo Santiago, Jeff Mu et al.ICML 2024 · 5 citations
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 458 citations
- Rethinking Shapley Value for Negative Interactions in Non-convex GamesWonjoon Chang, Myeongjin Lee, Jaesik ChoiICLR 2025
- Explanatory Model Monitoring to Understand the Effects of Feature Shifts on PerformanceThomas Decker, Alexander Koebler, Michael Lebacher, Ingo Thon et al.KDD 2024
