Mining the Minoria: Unknown, Under-represented, and Under-performing Minority Groups
Mohsen Dehghankar, Abolfazl Asudeh
摘要
Due to a variety of reasons, such as privacy, data in the wild often misses the grouping information required for identifying minorities. On the other hand, it is known that machine learning models are only as good as the data they are trained on and, hence, may underperform for the under-represented minority groups. The missing grouping information presents a dilemma for responsible data scientists who find themselves in an unknown-unknown situation, where not only do they not have access to the grouping attributes but do not also know what groups to consider. This paper is an attempt to address this dilemma. Specifically, we propose a minority mining problem, where we find vectors in the attribute space that reveal potential groups that are under-represented and under-performing. Technically speaking, we propose a geometric transformation of data into a dual space and use notions such as the arrangement of hyperplanes to design an efficient algorithm for the problem in lower dimensions. Generalizing our solution to the higher dimensions is cursed by dimensionality. Therefore, we propose a solution based on smart exploration of the search space for such cases. We conduct comprehensive experiments using real-world and synthetic datasets alongside the theoretical analysis. Our experiment results demonstrate the effectiveness of our proposed solutions in mining the unknown, under-represented, and under-performing minorities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Looking for Trouble: Analyzing Classifier Behavior via Pattern DivergenceEliana Pastor, Luca de Alfaro, Elena BaralisSIGMOD 2021 · 被引用 51 次
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 被引用 45 次
- Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning ModelsKi Hyun Tae, Steven Euijong WhangSIGMOD 2021 · 被引用 31 次
- Optimal Sublinear Sampling of Spanning Trees and Determinantal Point Processes via Average-Case Entropic IndependenceNima Anari, Yang P. Liu, Thuy-Duong VuongFOCS 2022 · 被引用 1 次
相关 Paper
- Exploratory Machine Learning with Unknown UnknownsPeng Zhao, Yu-Jie Zhang, Zhi-Hua ZhouAAAI 2021 · 被引用 29 次
- Boosting Test Performance with Importance Sampling-a Subpopulation PerspectiveHongyu Shen, Zhizhen ZhaoAAAI 2025
- Improving Subgroup Robustness via Data SelectionSaachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas 等NeurIPS 2024 · 被引用 17 次
- Tailoring Data Source Distributions for Fairness-aware Data IntegrationFatemeh Nargesian, Abolfazl Asudeh, H. V. JagadishVLDB 2021 · 被引用 51 次
- Identification of Systematic Errors of Image Classifiers on Rare SubgroupsJan Hendrik Metzen, Robin Hutmacher, N. Grace Hua, Valentyn Boreiko 等ICCV 2023 · 被引用 23 次
