Robustness to Spurious Correlations via Human Annotations
Megha Srivastava, Tatsunori B. Hashimoto, Percy Liang
摘要
The reliability of machine learning systems critically assumes that the associations between features and labels remain similar between training and test distributions. However, unmeasured variables, such as confounders, break this assumption---useful correlations between features and labels at training time can become useless or even harmful at test time. For example, high obesity is generally predictive for heart disease, but this relation may not hold for smokers who generally have lower rates of obesity and higher rates of heart disease. We present a framework for making models robust to spurious correlations by leveraging humans' common sense knowledge of causality. Specifically, we use human annotation to augment each training example with a potential unmeasured variable (i.e. an underweight patient with heart disease may be a smoker), reducing the problem to a covariate shift problem. We then introduce a new distributionally robust optimization objective over unmeasured variables (UV-DRO) to control the worst-case loss over possible test-time shifts. Empirically, we show improvements of 5-10% on a digit recognition task confounded by rotation, and 1.5-5% on the task of analyzing NYPD Police Stops confounded by location.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini 等NeurIPS 2020 · 被引用 731 次
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 被引用 454 次
- Robustness to Spurious Correlations in Text Classification via Automatically Generated CounterfactualsZhao Wang, Aron CulottaAAAI 2021 · 被引用 114 次
- ZIN: When and How to Learn Invariance Without Environment Partition?Yong Lin, Shengyu Zhu, Lu Tan, Peng CuiNeurIPS 2022 · 被引用 91 次
它引用的顶会 Paper1
相关 Paper
- Examining and Combating Spurious Features under Distribution ShiftChunting Zhou, Xuezhe Ma, Paul Michel, Graham NeubigICML 2021 · 被引用 78 次
- Mitigating Spurious Correlation via Distributionally Robust Learning with Hierarchical Ambiguity SetsSung Ho Jo, Seonghwi Kim, Minwoo ChaeICLR 2026 · 被引用 6 次
- Distributionally Robust Optimization with Probabilistic GroupSoumya Suvra Ghosal, Yixuan LiAAAI 2023 · 被引用 14 次
- Label-Efficient Group Robustness via Out-of-Distribution Concept CurationYiwei Yang, Anthony Z. Liu, Robert Wolfe, Aylin Caliskan 等CVPR 2024
- Distributionally Robust Classification for Multi-source Unsupervised Domain AdaptationSeonghwi Kim, Sungho Jo, Wooseok Ha, Minwoo ChaeICLR 2026 · 被引用 4 次
