Towards Explaining Distribution Shifts
Sean Kulinski, David I. Inouye
Abstract
A distribution shift can have fundamental consequences such as signaling a change in the operating environment or significantly reducing the accuracy of downstream models. Thus, understanding distribution shifts is critical for examining and hopefully mitigating the effect of such a shift. Most prior work focuses on merely detecting if a shift has occurred and assumes any detected shift can be understood and handled appropriately by a human operator. We hope to aid in these manual mitigation tasks by explaining the distribution shift using interpretable transportation maps from the original distribution to the shifted one. We derive our interpretable mappings from a relaxation of optimal transport, where the candidate mappings are restricted to a set of interpretable mappings. We then inspect multiple quintessential use-cases of distribution shift in real-world tabular, text, and image datasets to showcase how our explanatory mappings provide a better balance between detail and interpretability than baseline explanations by both visual inspection and our PercentExplained metric.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78182128-396c-41a2-9de7-91a2991d6c61Cited by top-tier papers13
- "Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsHaoran Zhang, Harvineet Singh, Marzyeh Ghassemi, Shalmali JoshiICML 2023 · 37 citations
- Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World EnvironmentsPaulius Rauba, Nabeel Seedat, Krzysztof Kacprzyk, Mihaela van der SchaarNeurIPS 2024 · 15 citations
- Counterfactual Fairness by Combining Factual and Counterfactual PredictionsZeyu Zhou, Tianci Liu, Ruqi Bai, Jing Gao et al.NeurIPS 2024 · 11 citations
- A hierarchical decomposition for explaining ML performance discrepanciesHarvineet Singh, Fan Xia, Adarsh Subbaswamy, Alexej Gossmann et al.NeurIPS 2024 · 9 citations
- GraphChain: Large Language Models for Large-scale Graph Analysis via Tool ChainingChunyu Wei, Wenji Hu, Xingjia Hao, Xin Wang et al.NeurIPS 2025 · 7 citations
Builds on5
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Optimal transport mapping via input convex neural networksAshok Vardhan Makkuva, Amirhossein Taghvaei, Sewoong Oh, Jason D. LeeICML 2020 · 254 citations
- Counterfactual Generative NetworksAxel Sauer, Andreas GeigerICLR 2021 · 145 citations
- Wasserstein-2 Generative NetworksAlexander Korotin, Vage Egiazarian, Arip Asadulaev, Alexander Safin et al.ICLR 2021 · 128 citations
- Feature Shift Detection: Localizing Which Features Have Shifted via Conditional Distribution TestsSean Kulinski, Saurabh Bagchi, David I. InouyeNeurIPS 2020 · 39 citations
Related papers
- Explanatory Model Monitoring to Understand the Effects of Feature Shifts on PerformanceThomas Decker, Alexander Koebler, Michael Lebacher, Ingo Thon et al.KDD 2024
- MOT: Masked Optimal Transport for Partial Domain AdaptationYou-Wei Luo, Chuan-Xian RenCVPR 2023
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell et al.ICCV 2021 · 141 citations
- Text-Transport: Toward Learning Causal Effects of Natural LanguageVictoria Lin, Louis-Philippe Morency, Eli Ben-MichaelEMNLP 2023 · 3 citations
- MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training ConflictsWeixin Liang, James ZouICLR 2022 · 103 citations
