Distribution Regression with Sliced Wasserstein Kernels
Dimitri Meunier, Massimiliano Pontil, Carlo Ciliberto
Abstract
The problem of learning functions over spaces of probabilities - or distribution regression - is gaining significant interest in the machine learning community. A key challenge behind this problem is to identify a suitable representation capturing all relevant properties of the underlying functional mapping. A principled approach to distribution regression is provided by kernel mean embeddings, which lifts kernel-induced similarity on the input domain at the probability level. This strategy effectively tackles the two-stage sampling nature of the problem, enabling one to derive estimators with strong statistical guarantees, such as universal consistency and excess risk bounds. However, kernel mean embeddings implicitly hinge on the maximum mean discrepancy (MMD), a metric on probabilities, which may fail to capture key geometrical relations between distributions. In contrast, optimal transport (OT) metrics, are potentially more appealing. In this work, we propose an OT-based estimator for distribution regression. We build on the Sliced Wasserstein distance to obtain an OT-based representation. We study the theoretical properties of a kernel ridge regression estimator based on such representation, for which we prove universal consistency and excess risk bounds. Preliminary experiments complement our theoretical findings by showing the effectiveness of the proposed approach and compare it with MMD-based estimators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6212790-8c37-4deb-9b7d-56b30bbfa033Cited by top-tier papers8
- Sliced-Wasserstein on Symmetric Positive Definite Matrices for M/EEG SignalsClément Bonet, Benoît Malézieux, Alain Rakotomamonjy, Lucas Drumetz et al.ICML 2023 · 28 citations
- Sliced-Wasserstein Estimation with Spherical Harmonics as Control VariatesRémi Leluc, Aymeric Dieuleveut, François Portier, Johan Segers et al.ICML 2024 · 9 citations
- On Statistical Learning Theory for Distributional InputsChristian Fiedler, Pierre-François Massiani, Friedrich Solowjow, Sebastian TrimpeICML 2024 · 3 citations
- Learning to Embed Distributions via Maximum Kernel EntropyOleksii Kachaiev, Stefano RecanatesiNeurIPS 2024 · 3 citations
- Flowing Datasets with Wasserstein over Wasserstein Gradient FlowsClément Bonet, Christophe Vauthier, Anna KorbaICML 2025
Builds on3
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 138 citations
- Statistical and Topological Properties of Sliced Probability DivergencesKimia Nadjahi, Alain Durmus, Lénaïc Chizat, Soheil Kolouri et al.NeurIPS 2020 · 115 citations
- The Advantage of Conditional Meta-Learning for Biased Regularization and Fine TuningGiulia Denevi, Massimiliano Pontil, Carlo CilibertoNeurIPS 2020 · 42 citations
Related papers
- Kernel Quantile Embeddings and Associated Probability MetricsMasha Naslidnyk, Siu Lun Chau, François-Xavier Briol, Krikamol MuandetICML 2025
- Statistical Optimal Transport posed as Learning Kernel EmbeddingJagarlapudi Saketha Nath, Pratik Kumar JawanpuriaNeurIPS 2020 · 18 citations
- Online Sinkhorn: Optimal Transport distances from sample streamsArthur Mensch, Gabriel PeyréNeurIPS 2020 · 35 citations
- A Universal Approximation Theorem of Deep Neural Networks for Expressing Probability DistributionsYulong Lu, Jianfeng LuNeurIPS 2020 · 146 citations
- Diffeomorphic Mesh Deformation via Efficient Optimal Transport for Cortical Surface ReconstructionThanh-Tung Le, Khai Nguyen, Shanlin Sun, Kun Han et al.ICLR 2024 · 9 citations
