Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test Data
Boris van Breugel, Nabeel Seedat, Fergus Imrie, Mihaela van der Schaar
Abstract
Evaluating the performance of machine learning models on diverse and underrepresented subgroups is essential for ensuring fairness and reliability in real-world applications. However, accurately assessing model performance becomes challenging due to two main issues: (1) a scarcity of test data, especially for small subgroups, and (2) possible distributional shifts in the model's deployment setting, which may not align with the available test data. In this work, we introduce 3S Testing, a deep generative modeling framework to facilitate model evaluation by generating synthetic test sets for small subgroups and simulating distributional shifts. Our experiments demonstrate that 3S Testing outperforms traditional baselines -- including real test data alone -- in estimating model performance on minority subgroups and under plausible distributional shifts. In addition, 3S offers intervals around its performance estimates, exhibiting superior coverage of the ground truth compared to existing approaches. Overall, these results raise the question of whether we need a paradigm shift away from limited real test data towards synthetic test data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a4e292f-d773-4392-84f4-ce3a5b6afb9cCited by top-tier papers6
- ClavaDDPM: Multi-relational Data Synthesis with Cluster-guided Diffusion ModelsWei Pang, Masoumeh Shafieinejad, Lucy Liu, Stephanie Hazlewood et al.NeurIPS 2024 · 39 citations
- TAI3: Testing Agent Integrity in Interpreting User IntentShiwei Feng, Xiangzhe Xu, Xuan Chen, Kaiyuan Zhang et al.NeurIPS 2025 · 9 citations
- Context-Aware Testing: A New Paradigm for Model Testing with Large Language ModelsPaulius Rauba, Nabeel Seedat, Max Ruiz Luyten, Mihaela van der SchaarNeurIPS 2024 · 8 citations
- Generative Conditional Distributions by Neural (Entropic) Optimal TransportBao Nguyen, Binh Nguyen, Hieu Trung Nguyen, Viet Anh NguyenICML 2024 · 2 citations
- Fill In The Gaps: Model Calibration and Generalization with Synthetic DataYang Ba, Michelle Mancenido, Rong PanEMNLP 2024 · 2 citations
Builds on18
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative ModelsAhmed M. Alaa, Boris van Breugel, Evgeny S. Saveliev, Mihaela van der SchaarICML 2022 · 287 citations
- DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative NetworksBoris van Breugel, Trent Kyono, Jeroen Berrevoets, Mihaela van der SchaarNeurIPS 2021 · 174 citations
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur et al.ICLR 2022 · 160 citations
Related papers
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras et al.NeurIPS 2025 · 9 citations
- Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic DataBoris van Breugel, Zhaozhi Qian, Mihaela van der SchaarICML 2023 · 45 citations
- Estimating and Explaining Model Performance When Both Covariates and Labels ShiftLingjiao Chen, Matei Zaharia, James Y. ZouNeurIPS 2022 · 34 citations
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 149 citations
- Improving Subgroup Robustness via Data SelectionSaachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas et al.NeurIPS 2024 · 17 citations
