Statistical Multicriteria Benchmarking via the GSD-Front
Christoph Jansen, Georg Schollmeyer, Julian Rodemann, Hannah Blocher, Thomas Augustin
Abstract
Given the vast number of classifiers that have been (and continue to be) proposed, reliable methods for comparing them are becoming increasingly important. The desire for reliability is broken down into three main aspects: (1) Comparisons should allow for different quality metrics simultaneously. (2) Comparisons should take into account the statistical uncertainty induced by the choice of benchmark suite. (3) The robustness of the comparisons under small deviations in the underlying assumptions should be verifiable. To address (1), we propose to compare classifiers using a generalized stochastic dominance ordering (GSD) and present the GSD-front as an information-efficient alternative to the classical Pareto-front. For (2), we propose a consistent statistical estimator for the GSD-front and construct a statistical test for whether a (potentially new) classifier lies in the GSD-front of a set of state-of-the-art classifiers. For (3), we relax our proposed test using techniques from robust statistics and imprecise probabilities. We illustrate our concepts on the benchmark suite PMLB and on the platform OpenML.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbdfde14-7e2e-4ad2-9aa0-f09a143309e9Cited by top-tier papers2
- FedGPS: Statistical Rectification Against Data Heterogeneity in Federated LearningZhiqin Yang, Yonggang Zhang, Chenxin Li, Yiu-ming Cheung et al.NeurIPS 2025 · 7 citations
- G2: Guided Generation for Enhanced Output Diversity in LLMsZhiwen Ruan, Yixia Li, Yefeng Liu, Yun Chen et al.EMNLP 2025
Builds on1
Related papers
- Nonparametric LLM Evaluation from Preference DataDennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan FeuerriegelICML 2026 · 5 citations
- Expanding the AI Evaluation Toolbox with Statistical ModelsDrew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K Rao et al.ICML 2026 · 4 citations
- Describing Subjective Experiment Consistency by p-Value P-P PlotJakub Nawala, Lucjan Janowski, Bogdan Cmiel, Krzysztof RusekACM MM 2020 · 9 citations
- Multivariate Stochastic Dominance via Optimal Transport and Applications to Models BenchmarkingGabriel Rioux, Apoorva Nitsure, Mattia Rigotti, Kristjan H. Greenewald et al.NeurIPS 2024 · 6 citations
- Center-Outward q-Dominance: A Sample-Computable Proxy for Strong Stochastic Dominance in Stochastic Multi-Objective OptimisationRobin van der Laag, Hao Wang, Thomas Bäck, Yingjie FanAAAI 2026
