MMD-Fuse: Learning and Combining Kernels for Two-Sample Testing Without Data Splitting
Felix Biggs, Antonin Schrab, Arthur Gretton
Abstract
We propose novel statistics which maximise the power of a two-sample test based on the Maximum Mean Discrepancy (MMD), by adapting over the set of kernels used in defining it. For finite sets, this reduces to combining (normalised) MMD values under each of these kernels via a weighted soft maximum. Exponential concentration bounds are proved for our proposed statistics under the null and alternative. We further show how these kernels can be chosen in a data-dependent but permutation-independent way, in a well-calibrated test, avoiding data splitting. This technique applies more broadly to general permutation-based MMD testing, and includes the use of deep kernels with features learnt using unsupervised models such as auto-encoders. We highlight the applicability of our MMD-FUSE test on both synthetic low-dimensional and real-world high-dimensional data, and compare its performance in terms of power against current state-of-the-art kernel tests.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ed7f439-b0f0-4181-bed0-382c40df400cCited by top-tier papers12
- Particle Semi-Implicit Variational InferenceJen Ning Lim, Adam M. JohansenNeurIPS 2024 · 13 citations
- DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence TestingZhijian Zhou, Xunye Tian, Liuhua Peng, Chao Lei et al.NeurIPS 2025 · 8 citations
- Neural-Kernel Conditional Mean EmbeddingsEiki Shimizu, Kenji Fukumizu, Dino SejdinovicICML 2024 · 6 citations
- Knowledge Distillation Detection for Open-weights ModelsQin Shi, Amber Yijia Zheng, Qifan Song, Raymond A. YehNeurIPS 2025 · 4 citations
- Anchor-based Maximum Discrepancy for Relative Similarity TestingZhijian Zhou, Liuhua Peng, Xunye Tian, Feng LiuNeurIPS 2025 · 2 citations
Builds on20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- Maximum Mean Discrepancy Test is Aware of Adversarial AttacksRuize Gao, Feng Liu, Jingfeng Zhang, Bo Han et al.ICML 2021 · 77 citations
- A permutation-free kernel two-sample testShubhanshu Shekhar, Ilmun Kim, Aaditya RamdasNeurIPS 2022 · 40 citations
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang et al.ICML 2020 · 213 citations
- Kernel-based Maximum-of-difference Test for Two-sample ComparisonDan Pu, Tianyi Zhu, Yao Yan, Wei LanICML 2026
- Neural Tangent Kernel Maximum Mean DiscrepancyXiuyuan Cheng, Yao XieNeurIPS 2021 · 26 citations
