Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking
Gabriel Rioux, Apoorva Nitsure, Mattia Rigotti, Kristjan H. Greenewald, Youssef Mroueh
摘要
Stochastic dominance is an important concept in probability theory, econometrics and social choice theory for robustly modeling agents' preferences between random outcomes. While many works have been dedicated to the univariate case, little has been done in the multivariate scenario, wherein an agent has to decide between different multivariate outcomes. By exploiting a characterization of multivariate first stochastic dominance in terms of couplings, we introduce a statistic that assesses multivariate almost stochastic dominance under the framework of Optimal Transport with a smooth cost. Further, we introduce an entropic regularization of this statistic, and establish a central limit theorem (CLT) and consistency of the bootstrap procedure for the empirical statistic. Armed with this CLT, we propose a hypothesis testing framework as well as an efficient implementation using the Sinkhorn algorithm. We showcase our method in comparing and benchmarking Large Language Models that are evaluated on multiple metrics. Our multivariate stochastic dominance test allows us to capture the dependencies between the metrics in order to make an informed and statistically significant decision on the relative performance of the models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Imitation Beyond Expectation Using Pluralistic Stochastic DominanceAli Farajzadeh, Danyal Saeed, Syed M. Abbas, Rushit N. Shah 等NeurIPS 2025 · 被引用 2 次
- Center-Outward q-Dominance: A Sample-Computable Proxy for Strong Stochastic Dominance in Stochastic Multi-Objective OptimisationRobin van der Laag, Hao Wang, Thomas Bäck, Yingjie FanAAAI 2026
它引用的顶会 Paper4
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionDongfu Jiang, Xiang Ren, Bill Yuchen LinACL 2023 · 被引用 95 次
- Smooth p-Wasserstein Distance: Structure, Empirical Approximation, and Statistical ApplicationsSloan Nietert, Ziv Goldfeld, Kengo KatoICML 2021 · 被引用 39 次
- Inherent Trade-Offs between Diversity and Stability in Multi-Task BenchmarksGuanhua Zhang, Moritz HardtICML 2024 · 被引用 22 次
相关 Paper
- Risk Aware Benchmarking of Large Language ModelsApoorva Nitsure, Youssef Mroueh, Mattia Rigotti, Kristjan H. Greenewald 等ICML 2024 · 被引用 3 次
- Sinkhorn Treatment Effects: A Causal Optimal Transport MeasureMedha Agarwal, Alex LuedtkeICML 2026
- Bisimulation Metrics are Optimal Transport Distances, and Can be Computed EfficientlySergio Calo, Anders Jonsson, Gergely Neu, Ludovic Schwartz 等NeurIPS 2024 · 被引用 9 次
- Online Sinkhorn: Optimal Transport distances from sample streamsArthur Mensch, Gabriel PeyréNeurIPS 2020 · 被引用 35 次
- Learning with Stochastic OrdersCarles Domingo-Enrich, Yair Schiff, Youssef MrouehICLR 2023
