Learning Kernel Tests Without Data Splitting
Jonas M. Kübler, Wittawat Jitkrittum, Bernhard Schölkopf, Krikamol Muandet
Abstract
Modern large-scale kernel-based tests such as maximum mean discrepancy (MMD) and kernelized Stein discrepancy (KSD) optimize kernel hyperparameters on a held-out sample via data splitting to obtain the most powerful test statistics. While data splitting results in a tractable null distribution, it suffers from a reduction in test power due to smaller test sample size. Inspired by the selective inference framework, we propose an approach that enables learning the hyperparameters and testing on the full sample without data splitting. Our approach can correctly calibrate the test in the presence of such dependency, and yield a test threshold in closed form. At the same significance level, our approach's test power is empirically larger than that of the data-splitting approach, regardless of its split proportion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 392040c4-526d-48f6-82c2-e55539633290Cited by top-tier papers12
- MMD-Fuse: Learning and Combining Kernels for Two-Sample Testing Without Data SplittingFelix Biggs, Antonin Schrab, Arthur GrettonNeurIPS 2023 · 49 citations
- Efficient Aggregated Kernel Tests using Incomplete -statisticsAntonin Schrab, Ilmun Kim, Benjamin Guedj, Arthur GrettonNeurIPS 2022 · 42 citations
- Meta Two-Sample Testing: Learning Kernels for Testing with Limited DataFeng Liu, Wenkai Xu, Jie Lu, Danica J. SutherlandNeurIPS 2021 · 30 citations
- AutoML Two-Sample TestJonas M. Kübler, Vincent Stimper, Simon Buchholz, Krikamol Muandet et al.NeurIPS 2022 · 29 citations
- Neural Tangent Kernel Maximum Mean DiscrepancyXiuyuan Cheng, Yao XieNeurIPS 2021 · 26 citations
Builds on1
Related papers
- A permutation-free kernel two-sample testShubhanshu Shekhar, Ilmun Kim, Aaditya RamdasNeurIPS 2022 · 40 citations
- KSD Aggregated Goodness-of-fit TestAntonin Schrab, Benjamin Guedj, Arthur GrettonNeurIPS 2022 · 26 citations
- Anchor-based Maximum Discrepancy for Relative Similarity TestingZhijian Zhou, Liuhua Peng, Xunye Tian, Feng LiuNeurIPS 2025 · 2 citations
- Sliced Kernelized Stein DiscrepancyWenbo Gong, Yingzhen Li, José Miguel Hernández-LobatoICLR 2021 · 14 citations
- The Polynomial Stein Discrepancy for Assessing Moment ConvergenceNarayan Srinivasan, Matthew Sutton, Christopher C. Drovandi, Leah F. SouthICML 2025
