Data-IQ: Characterizing subgroups with heterogeneous outcomes in tabular data
Nabeel Seedat, Jonathan Crabbé, Ioana Bica, Mihaela van der Schaar
Abstract
High model performance, on average, can hide that models may systematically underperform on subgroups of the data. We consider the tabular setting, which surfaces the unique issue of outcome heterogeneity -this is prevalent in areas such as healthcare, where patients with similar features can have different outcomes, thus making reliable predictions challenging. To tackle this, we propose Data-IQ, a framework to systematically stratify examples into subgroups with respect to their outcomes. We do this by analyzing the behavior of individual examples during training, based on their predictive confidence and, importantly, the aleatoric (data) uncertainty. Capturing the aleatoric uncertainty permits a principled characterization and then subsequent stratification of data examples into three distinct subgroups (Easy, Ambiguous, Hard). We experimentally demonstrate the benefits of Data-IQ on four real-world medical datasets. We show that Data-IQ's characterization of examples is most robust to variation across similarly performant (yet different) models, compared to baselines. Since Data-IQ can be used with any ML model (including neural networks, gradient boosting etc.), this property ensures consistency of data characterization, while allowing flexible model selection. Taking this a step further, we demonstrate that the subgroups enable us to construct new approaches to both feature acquisition and dataset selection. Furthermore, we highlight how the subgroups can inform reliable model usage, noting the significant impact of the Ambiguous subgroup on model generalization. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimesNabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der SchaarICML 2024 · 61 citations
- Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test DataBoris van Breugel, Nabeel Seedat, Fergus Imrie, Mihaela van der SchaarNeurIPS 2023 · 51 citations
- Dissecting Sample Hardness: A Fine-Grained Analysis of Hardness Characterization Methods for Data-Centric AINabeel Seedat, Fergus Imrie, Mihaela van der SchaarICLR 2024 · 16 citations
- TRIAGE: Characterizing and auditing training data for improved regressionNabeel Seedat, Jonathan Crabbé, Zhaozhi Qian, Mihaela van der SchaarNeurIPS 2023 · 8 citations
- Breaking the Barrier of Hard Samples: A Data-Centric Approach to Synthetic Data for Medical TasksMaynara Donato de Souza, Cleber ZanchettinICML 2025
Builds on11
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular DomainJinsung Yoon, Yao Zhang, James Jordon, Mihaela van der SchaarNeurIPS 2020 · 370 citations
Related papers
- Data-Driven Subgroup Identification for Linear RegressionZachary Izzo, Ruishan Liu, James ZouICML 2023 · 7 citations
- Robust Recursive Partitioning for Heterogeneous Treatment Effects with Uncertainty QuantificationHyun-Suk Lee, Yao Zhang, William R. Zame, Cong Shen et al.NeurIPS 2020 · 21 citations
- Unveiling the Role of Data Uncertainty in Tabular Deep LearningNikolay Kartashev, Ivan Rubachev, Artem BabenkoICML 2026 · 1 citation
- Data-SUITE: Data-centric identification of in-distribution incongruous examplesNabeel Seedat, Jonathan Crabbé, Mihaela van der SchaarICML 2022 · 16 citations
- Improving Subgroup Robustness via Data SelectionSaachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas et al.NeurIPS 2024 · 17 citations
