Representation Matters: Assessing the Importance of Subgroup Allocations in Training Data
Esther Rolf, Theodora T. Worledge, Benjamin Recht, Michael I. Jordan
摘要
Collecting more diverse and representative training data is often touted as a remedy for the disparate performance of machine learning predictors across subpopulations. However, a precise framework for understanding how dataset properties like diversity affect learning outcomes is largely lacking. By casting data collection as part of the learning process, we demonstrate that diverse representation in training data is key not only to increasing subgroup performances, but also to achieving population level objectives. Our analysis and experiments describe how dataset compositions influence performance and provide constructive results for using trends in existing data, alongside domain knowledge, to help guide intentional, objective-aware dataset design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Active Sampling for Min-Max FairnessJacob D. Abernethy, Pranjal Awasthi, Matthäus Kleindessner, Jamie Morgenstern 等ICML 2022 · 被引用 57 次
- The Illusion of Artificial InclusionWilliam Agnew, A. Stevie Bergman, Jennifer Chien, Mark Díaz 等CHI 2024 · 被引用 57 次
- Scaling Laws for the Value of Individual Data Points in Machine LearningIan Connick Covert, Wenlong Ji, Tatsunori Hashimoto, James ZouICML 2024 · 被引用 12 次
- Fairness in model-sharing gamesKate Donahue, Jon M. KleinbergWWW 2023 · 被引用 12 次
- Mind the Graph When Balancing Data for Fairness or RobustnessJessica Schrouff, Alexis Bellot, Amal Rannen-Triki, Alan Malek 等NeurIPS 2024 · 被引用 10 次
它引用的顶会 Paper2
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification ProblemsNimit Sharad Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu 等NeurIPS 2020 · 被引用 316 次
相关 Paper
- Fair Classification with Partial Feedback: An Exploration-Based Data Collection ApproachVijay Keswani, Anay Mehrotra, L. Elisa CelisICML 2024 · 被引用 3 次
- Diverse Prototypical Ensembles Improve Robustness to Subpopulation ShiftMinh Nguyen Nhat To, Paul F. R. Wilson, Viet Nguyen, Mohamed Harmanani 等ICML 2025
- On Harmonizing Implicit SubpopulationsFeng Hong, Jiangchao Yao, Yueming Lyu, Zhihan Zhou 等ICLR 2024 · 被引用 8 次
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 被引用 149 次
- Multigroup RobustnessLunjia Hu, Charlotte Peale, Judy Hanwen ShenICML 2024 · 被引用 2 次
