Data-SUITE: Data-centric identification of in-distribution incongruous examples
Nabeel Seedat, Jonathan Crabbé, Mihaela van der Schaar
摘要
Systematic quantification of data quality is critical for consistent model performance. Prior works have focused on out-of-distribution data. Instead, we tackle an understudied yet equally important problem of characterizing incongruous regions of in-distribution (ID) data, which may arise from feature space heterogeneity. To this end, we propose a paradigm shift with Data-SUITE: a data-centric AI framework to identify these regions, independent of a task-specific model. Data-SUITE leverages copula modeling, representation learning, and conformal prediction to build feature-wise confidence interval estimators based on a set of training instances. These estimators can be used to evaluate the congruence of test instances with respect to the training set, to answer two practically useful questions: (1) which test instances will be reliably predicted by a model trained with the training instances? and (2) can we identify incongruous regions of the feature space so that data owners understand the data's limitations or guide future data collection? We empirically validate Data-SUITE's performance and coverage guarantees and demonstrate on cross-site medical data, biased data, and data with concept drift, that Data-SUITE best identifies ID regions where a downstream model may be reliable (independent of said model). We also illustrate how these identified regions can provide insights into datasets and highlight their limitations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test DataBoris van Breugel, Nabeel Seedat, Fergus Imrie, Mihaela van der SchaarNeurIPS 2023 · 被引用 51 次
- Data-IQ: Characterizing subgroups with heterogeneous outcomes in tabular dataNabeel Seedat, Jonathan Crabbé, Ioana Bica, Mihaela van der SchaarNeurIPS 2022 · 被引用 40 次
- Investigating Generalizability of Speech-based Suicidal Ideation Detection Using Mobile PhonesArvind Pillai, Subigya Kumar Nepal, Weichen Wang, Matthew Nemesure 等UbiComp 2024 · 被引用 26 次
- Dissecting Sample Hardness: A Fine-Grained Analysis of Hardness Characterization Methods for Data-Centric AINabeel Seedat, Fergus Imrie, Mihaela van der SchaarICLR 2024 · 被引用 16 次
- TRIAGE: Characterizing and auditing training data for improved regressionNabeel Seedat, Jonathan Crabbé, Zhaozhi Qian, Mihaela van der SchaarNeurIPS 2023 · 被引用 8 次
它引用的顶会 Paper6
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Understanding Failures in Out-of-Distribution Detection with Deep Generative ModelsLily H. Zhang, Mark Goldstein, Rajesh RanganathICML 2021 · 被引用 129 次
- Discriminative Jackknife: Quantifying Uncertainty in Deep Learning via Higher-Order Influence FunctionsAhmed M. Alaa, Mihaela van der SchaarICML 2020 · 被引用 59 次
- Reliable and Trustworthy Machine Learning for Health Using Dataset Shift DetectionChunjong Park, Anas Awadalla, Tadayoshi Kohno, Shwetak N. PatelNeurIPS 2021 · 被引用 50 次
- Unlabelled Data Improves Bayesian Uncertainty Calibration under Covariate ShiftAlex J. Chan, Ahmed M. Alaa, Zhaozhi Qian, Mihaela van der SchaarICML 2020 · 被引用 42 次
相关 Paper
- Adaptive Conformal Prediction Intervals for Invariant LearningShuxin Liang, Yihan Xiao, Linglong Kong, Wenlu TangKDD 2025
- Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)Drew Prinster, Samuel Don Stanton, Anqi Liu, Suchi SariaICML 2024 · 被引用 20 次
- COMPASS: Robust Feature Conformal Prediction for Medical Segmentation MetricsMatt Y. Cheung, Ashok Veeraraghavan, Guha BalakrishnanICLR 2026 · 被引用 5 次
- Learning Prediction Intervals for Model PerformanceBenjamin Elder, Matthew Arnold, Anupama Murthi, Jirí NavrátilAAAI 2021 · 被引用 14 次
- Relational Conformal Prediction for Correlated Time SeriesAndrea Cini, Alexander Jenkins, Danilo P. Mandic, Cesare Alippi 等ICML 2025
