Quantifying and Mitigating the Impact of Label Errors on Model Disparity Metrics
Julius Adebayo, Melissa Hall, Bowen Yu, Bobbie Chern
摘要
Errors in labels obtained via human annotation adversely affect a model's performance. Existing approaches propose ways to mitigate the effect of label error on a model's downstream accuracy, yet little is known about its impact on a model's disparity metrics. Here we study the effect of label error on a model's disparity metrics. We empirically characterize how varying levels of label error, in both training and test data, affect these disparity metrics. We find that group calibration and other metrics are sensitive to train-time and test-time label error -- particularly for minority groups. This disparate effect persists even for models trained with noise-aware algorithms. To mitigate the impact of training-time label error, we present an approach to estimate the influence of a training input's label on a model's group disparity metric. We empirically assess the proposed approach on a variety of datasets and find significant improvement, compared to alternative approaches, in identifying training inputs that improve a model's disparity metric. We complement the approach with an automatic relabel-and-finetune scheme that produces updated models with, provably, improved group calibration error.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and ResolutionMostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek 等NeurIPS 2023 · 被引用 303 次
- Improving Subgroup Robustness via Data SelectionSaachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas 等NeurIPS 2024 · 被引用 17 次
- Automating Data Annotation under Strategic Human Agents: Risks and Potential SolutionsTian Xie, Xueru ZhangNeurIPS 2024 · 被引用 12 次
- Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation ModelsTianjian Li, Haoran Xu, Philipp Koehn, Daniel Khashabi 等ICLR 2024 · 被引用 6 次
- On Robustness of Linear Classifiers to Targeted Data PoisoningNakshatra Gupta, Sumanth Prabhu S, Supratik Chakraborty, R. VenkateshAAAI 2026
它引用的顶会 Paper21
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano 等ICML 2020 · 被引用 547 次
相关 Paper
- Estimating Structural Disparities for Face ModelsShervin Ardeshir, Cristina Segalin, Nathan KallusCVPR 2022 · 被引用 2 次
- Enhancing Robustness of Last Layer Two-Stage Fair Model CorrectionsNathan Stromberg, Rohan Ayyagari, Sanmi Koyejo, Richard Nock 等NeurIPS 2024 · 被引用 3 次
- The Group Robustness is in the Details: Revisiting Finetuning under Spurious CorrelationsTyler LaBonte, John C. Hill, Xinchen Zhang, Vidya Muthukumar 等NeurIPS 2024 · 被引用 8 次
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 被引用 436 次
- On Group Sufficiency Under Label BiasHaoran Zhang, Olawale Salaudeen, Marzyeh GhassemiNeurIPS 2025 · 被引用 2 次
