Mandoline: Model Evaluation under Distribution Shift
Mayee F. Chen, Karan Goel, Nimit Sharad Sohoni, Fait Poms, Kayvon Fatahalian, Christopher Ré
摘要
Machine learning models are often deployed in different settings than they were trained and validated on, posing a challenge to practitioners who wish to predict how well the deployed model will perform on a target distribution. If an unlabeled sample from the target distribution is available, along with a labeled sample from a possibly different source distribution, standard approaches such as importance weighting can be applied to estimate performance on the target. However, importance weighting struggles when the source and target distributions have non-overlapping support or are high-dimensional. Taking inspiration from fields such as epidemiology and polling, we develop Mandoline, a new evaluation framework that mitigates these issues. Our key insight is that practitioners may have prior knowledge about the ways in which the distribution shifts, which we can use to better guide the importance weighting procedure. Specifically, users write simple"slicing functions"- noisy, potentially correlated binary functions intended to capture possible axes of distribution shift - to compute reweighted performance estimates. We further describe a density ratio estimation framework for the slices and show how its estimation error scales with slice quality and dataset size. Empirical validation on NLP and vision tasks shows that Mandoline can estimate performance on the target distribution up to 3x more accurately compared to standard baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck 等ICLR 2022 · 被引用 178 次
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur 等ICLR 2022 · 被引用 160 次
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell 等ICCV 2021 · 被引用 141 次
- Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution ShiftChristina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 被引用 120 次
- Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training EnsemblesJiefeng Chen, Frederick Liu, Besim Avci, Xi Wu 等NeurIPS 2021 · 被引用 79 次
它引用的顶会 Paper8
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Rethinking Importance Weighting for Deep Learning under Distribution ShiftTongtong Fang, Nan Lu, Gang Niu, Masashi SugiyamaNeurIPS 2020 · 被引用 179 次
- Telescoping Density-Ratio EstimationBenjamin Rhodes, Kai Xu, Michael U. GutmannNeurIPS 2020 · 被引用 148 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
相关 Paper
- Understanding new tasks through the lens of training data via exponential tiltingSubha Maity, Mikhail Yurochkin, Moulinath Banerjee, Yuekai SunICLR 2023 · 被引用 1 次
- ReTaSA: A Nonparametric Functional Estimation Approach for Addressing Continuous Target ShiftHwanwoo Kim, Xin Zhang, Jiwei Zhao, Qinglong TianICLR 2024 · 被引用 3 次
- Estimating Generalization under Distribution Shifts via Domain-Invariant RepresentationsChing-Yao Chuang, Antonio Torralba, Stefanie JegelkaICML 2020 · 被引用 72 次
- IW-GAE: Importance weighted group accuracy estimation for improved calibration and model selection in unsupervised domain adaptationTaejong Joo, Diego KlabjanICML 2024 · 被引用 1 次
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 被引用 60 次
