Conformal C2ST: Turning weak classifiers into strong two-sample tests
Vansh Bansal, Tianyu Chen, James Scott
Abstract
The two-sample testing problem, a fundamental task in statistics and machine learning, seeks to determine whether two sets of samples, drawn from underlying distributions p and q, are in fact identically distributed (i.e. whether p = q). A popular and intuitive approach is the classifier two-sample test (C2ST), where a classifier is trained to distinguish between samples from p and q. Yet despite simplicity of the C2ST, its reliability hinges on access to a near-Bayes-optimal classifier, a requirement that is rarely met and difficult to verify. This raises a major open question: can a weak classifier still be useful for two-sample testing? We show that the answer is a definitive yes. Building on the work of Hu & Lei (2024), we analyze two conformal variants of the C2ST that convert the scores from any trained classifier-even if weak, biased, or overfit-into exact, finite-sample p-values. We establish two key theoretical properties of the conformal C2ST: (i) finite-sample Type-I error control, and (ii) non-trivial power that degrades gently in tandem with the error of the trained classifier. The upshot is that even poorly performing classifiers can yield powerful and reliable two-sample tests. This general framework finds a powerful application in Bayesian inference, particularly for validating Neural Posterior Estimation (NPE) models, where the task of comparing a learned posterior approximation q(θ | y) to the true posterior p(θ | y) can be framed as a twosample test. Empirically, the Conformal C2ST outperforms classical discriminative tests across a wide range of benchmarks for this task. Our results establish the conformal C2ST as a practical, theoretically grounded diagnostic tool. The code is available at https://github.com/ TianyuCodings/conformal_c2st .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Flow Matching for Scalable Simulation-Based InferenceJonas Wildberger, Maximilian Dax, Simon Buchholz, Stephen R. Green et al.NeurIPS 2023 · 153 citations
- All-in-one simulation-based inferenceManuel Glöckler, Michael Deistler, Christian Dietrich Weilbach, Frank Wood et al.ICML 2024 · 74 citations
- Sampling-Based Accuracy Testing of Posterior Estimators for General InferencePablo Lemos, Adam Coogan, Yashar Hezaveh, Laurence Perreault LevasseurICML 2023 · 66 citations
- MMD-Fuse: Learning and Combining Kernels for Two-Sample Testing Without Data SplittingFelix Biggs, Antonin Schrab, Arthur GrettonNeurIPS 2023 · 49 citations
Related papers
- L-C2ST: Local Diagnostics for Posterior Approximations in Simulation-Based InferenceJulia Linhart, Alexandre Gramfort, Pedro RodriguesNeurIPS 2023 · 25 citations
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang et al.ICML 2020 · 213 citations
- CoLT: The conditional localization test for assessing the accuracy of neural posterior estimatesTianyu Chen, Vansh Bansal, James G. ScottNeurIPS 2025
- Conditional Testing based on Localized Conformal p-valuesXiaoyang Wu, Lin Lu, Zhaojun Wang, Changliang ZouICLR 2025
- Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary ClassificationTakashi Ishida, Ikko Yamane, Nontawat Charoenphakdee, Gang Niu et al.ICLR 2023 · 3 citations
