Mandoline: Model Evaluation under Distribution Shift
Mayee F. Chen, Karan Goel, Nimit Sharad Sohoni, Fait Poms, Kayvon Fatahalian, Christopher Ré
Abstract
Machine learning models are often deployed in different settings than they were trained and validated on, posing a challenge to practitioners who wish to predict how well the deployed model will perform on a target distribution. If an unlabeled sample from the target distribution is available, along with a labeled sample from a possibly different source distribution, standard approaches such as importance weighting can be applied to estimate performance on the target. However, importance weighting struggles when the source and target distributions have non-overlapping support or are high-dimensional. Taking inspiration from fields such as epidemiology and polling, we develop Mandoline, a new evaluation framework that mitigates these issues. Our key insight is that practitioners may have prior knowledge about the ways in which the distribution shifts, which we can use to better guide the importance weighting procedure. Specifically, users write simple"slicing functions"- noisy, potentially correlated binary functions intended to capture possible axes of distribution shift - to compute reweighted performance estimates. We further describe a density ratio estimation framework for the slices and show how its estimation error scales with slice quality and dataset size. Empirical validation on NLP and vision tasks shows that Mandoline can estimate performance on the target distribution up to 3x more accurately compared to standard baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e551337e-4ded-4c37-8d1d-bd27bb7ded20Cited by top-tier papers30
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck et al.ICLR 2022 · 178 citations
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur et al.ICLR 2022 · 160 citations
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell et al.ICCV 2021 · 141 citations
- Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution ShiftChristina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 120 citations
- Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training EnsemblesJiefeng Chen, Frederick Liu, Besim Avci, Xi Wu et al.NeurIPS 2021 · 79 citations
Builds on8
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Rethinking Importance Weighting for Deep Learning under Distribution ShiftTongtong Fang, Nan Lu, Gang Niu, Masashi SugiyamaNeurIPS 2020 · 179 citations
- Telescoping Density-Ratio EstimationBenjamin Rhodes, Kai Xu, Michael U. GutmannNeurIPS 2020 · 148 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- Understanding new tasks through the lens of training data via exponential tiltingSubha Maity, Mikhail Yurochkin, Moulinath Banerjee, Yuekai SunICLR 2023 · 1 citation
- ReTaSA: A Nonparametric Functional Estimation Approach for Addressing Continuous Target ShiftHwanwoo Kim, Xin Zhang, Jiwei Zhao, Qinglong TianICLR 2024 · 3 citations
- Estimating Generalization under Distribution Shifts via Domain-Invariant RepresentationsChing-Yao Chuang, Antonio Torralba, Stefanie JegelkaICML 2020 · 72 citations
- IW-GAE: Importance weighted group accuracy estimation for improved calibration and model selection in unsupervised domain adaptationTaejong Joo, Diego KlabjanICML 2024 · 1 citation
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 60 citations
