TRIAGE: Characterizing and auditing training data for improved regression
Nabeel Seedat, Jonathan Crabbé, Zhaozhi Qian, Mihaela van der Schaar
Abstract
Data quality is crucial for robust machine learning algorithms, with the recent interest in data-centric AI emphasizing the importance of training data characterization. However, current data characterization methods are largely focused on classification settings, with regression settings largely understudied. To address this, we introduce TRIAGE, a novel data characterization framework tailored to regression tasks and compatible with a broad class of regressors. TRIAGE utilizes conformal predictive distributions to provide a model-agnostic scoring method, the TRIAGE score. We operationalize the score to analyze individual samples' training dynamics and characterize samples as under-, over-, or well-estimated by the model. We show that TRIAGE's characterization is consistent and highlight its utility to improve performance via data sculpting/filtering, in multiple regression settings. Additionally, beyond sample level, we show TRIAGE enables new approaches to dataset selection and feature acquisition. Overall, TRIAGE highlights the value unlocked by data characterization in real-world regression applications. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimesNabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der SchaarICML 2024 · 61 citations
- Dissecting Sample Hardness: A Fine-Grained Analysis of Hardness Characterization Methods for Data-Centric AINabeel Seedat, Fergus Imrie, Mihaela van der SchaarICLR 2024 · 16 citations
- Accelerating Feature Conformal Prediction via Taylor ApproximationZihao Tang, Boyuan Wang, Chuan Wen, Jiaye TengNeurIPS 2025 · 1 citation
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema MatchingNabeel Seedat, Mihaela van der SchaarICML 2025
Builds on10
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- Estimating Example Difficulty using Variance of GradientsChirag Agarwal, Daniel D'souza, Sara HookerCVPR 2022 · 57 citations
- Characterizing Datapoints via Second-Split ForgettingPratyush Maini, Saurabh Garg, Zachary C. Lipton, J. Zico KolterNeurIPS 2022 · 48 citations
Related papers
- Data-SUITE: Data-centric identification of in-distribution incongruous examplesNabeel Seedat, Jonathan Crabbé, Mihaela van der SchaarICML 2022 · 16 citations
- Data-IQ: Characterizing subgroups with heterogeneous outcomes in tabular dataNabeel Seedat, Jonathan Crabbé, Ioana Bica, Mihaela van der SchaarNeurIPS 2022 · 40 citations
- Utility-Directed Conformal Prediction: A Decision-Aware Framework for Actionable Uncertainty QuantificationSantiago Cortes-Gomez, Carlos Miguel Patiño, Yewon Byun, Steven Wu et al.ICLR 2025
- Robust Conformal Outlier Detection under Contaminated Reference DataMeshi Bashari, Matteo Sesia, Yaniv RomanoICML 2025
- Fast Conformal Prediction Using Conditional Interquantile IntervalsNaixin Guo, Rui Luo, Zhixin ZhouAAAI 2026 · 4 citations
