Mining the Minoria: Unknown, Under-represented, and Under-performing Minority Groups
Mohsen Dehghankar, Abolfazl Asudeh
Abstract
Due to a variety of reasons, such as privacy, data in the wild often misses the grouping information required for identifying minorities. On the other hand, it is known that machine learning models are only as good as the data they are trained on and, hence, may underperform for the under-represented minority groups. The missing grouping information presents a dilemma for responsible data scientists who find themselves in an unknown-unknown situation, where not only do they not have access to the grouping attributes but do not also know what groups to consider. This paper is an attempt to address this dilemma. Specifically, we propose a minority mining problem, where we find vectors in the attribute space that reveal potential groups that are under-represented and under-performing. Technically speaking, we propose a geometric transformation of data into a dual space and use notions such as the arrangement of hyperplanes to design an efficient algorithm for the problem in lower dimensions. Generalizing our solution to the higher dimensions is cursed by dimensionality. Therefore, we propose a solution based on smart exploration of the search space for such cases. We conduct comprehensive experiments using real-world and synthetic datasets alongside the theoretical analysis. Our experiment results demonstrate the effectiveness of our proposed solutions in mining the unknown, under-represented, and under-performing minorities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0a49f00-b411-473e-a2b0-b4820ec4fa13Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Looking for Trouble: Analyzing Classifier Behavior via Pattern DivergenceEliana Pastor, Luca de Alfaro, Elena BaralisSIGMOD 2021 · 51 citations
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 45 citations
- Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning ModelsKi Hyun Tae, Steven Euijong WhangSIGMOD 2021 · 31 citations
- Optimal Sublinear Sampling of Spanning Trees and Determinantal Point Processes via Average-Case Entropic IndependenceNima Anari, Yang P. Liu, Thuy-Duong VuongFOCS 2022 · 1 citation
Related papers
- Exploratory Machine Learning with Unknown UnknownsPeng Zhao, Yu-Jie Zhang, Zhi-Hua ZhouAAAI 2021 · 29 citations
- Boosting Test Performance with Importance Sampling-a Subpopulation PerspectiveHongyu Shen, Zhizhen ZhaoAAAI 2025
- Improving Subgroup Robustness via Data SelectionSaachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas et al.NeurIPS 2024 · 17 citations
- Tailoring Data Source Distributions for Fairness-aware Data IntegrationFatemeh Nargesian, Abolfazl Asudeh, H. V. JagadishVLDB 2021 · 51 citations
- Identification of Systematic Errors of Image Classifiers on Rare SubgroupsJan Hendrik Metzen, Robin Hutmacher, N. Grace Hua, Valentyn Boreiko et al.ICCV 2023 · 23 citations
