Meta Two-Sample Testing: Learning Kernels for Testing with Limited Data
Feng Liu, Wenkai Xu, Jie Lu, Danica J. Sutherland
Abstract
Modern kernel-based two-sample tests have shown great success in distinguishing complex, high-dimensional distributions with appropriate learned kernels. Previous work has demonstrated that this kernel learning procedure succeeds, assuming a considerable number of observed samples from each distribution. In realistic scenarios with very limited numbers of data samples, however, it can be challenging to identify a kernel powerful enough to distinguish complex distributions. We address this issue by introducing the problem of meta two-sample testing (M2ST), which aims to exploit (abundant) auxiliary data on related tasks to find an algorithm that can quickly identify a powerful test on new target tasks. We propose two specific algorithms for this task: a generic scheme which improves over baselines and a more tailored approach which performs even better. We provide both theoretical justification and empirical evidence that our proposed meta-testing schemes out-perform learning kernel-based tests directly from scarce observations, and identify when such schemes will be successful.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b904da1-1dbc-4979-aa6a-22f7bd3dcb87Cited by top-tier papers16
- MMD-Fuse: Learning and Combining Kernels for Two-Sample Testing Without Data SplittingFelix Biggs, Antonin Schrab, Arthur GrettonNeurIPS 2023 · 49 citations
- AutoML Two-Sample TestJonas M. Kübler, Vincent Stimper, Simon Buchholz, Krikamol Muandet et al.NeurIPS 2022 · 29 citations
- Bilateral Dependency Optimization: Defending Against Model-inversion AttacksXiong Peng, Feng Liu, Jingfeng Zhang, Long Lan et al.KDD 2022 · 20 citations
- Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean DiscrepancyShuhai Zhang, Yiliao Song, Jiahao Yang, Yuanqing Li et al.ICLR 2024 · 19 citations
- DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence TestingZhijian Zhou, Xunye Tian, Liuhua Peng, Chao Lei et al.NeurIPS 2025 · 8 citations
Builds on12
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang et al.ICML 2020 · 213 citations
- Attribute Propagation Network for Graph Zero-Shot LearningLu Liu, Tianyi Zhou, Guodong Long, Jing Jiang et al.AAAI 2020 · 85 citations
- Learning Bounds for Open-Set LearningZhen Fang, Jie Lu, Anjin Liu, Feng Liu et al.ICML 2021 · 67 citations
- Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation LearningLuca Oneto, Michele Donini, Giulia Luise, Carlo Ciliberto et al.NeurIPS 2020 · 56 citations
- How Does the Combined Risk Affect the Performance of Unsupervised Domain Adaptation Approaches?Zhong Li, Zhen Fang, Feng Liu, Jie Lu et al.AAAI 2021 · 56 citations
Related papers
- Effective Meta-Regularization by Kernelized Proximal RegularizationWeisen Jiang, James T. Kwok, Yu ZhangNeurIPS 2021 · 9 citations
- Adversarial Task Up-sampling for Meta-learningYichen Wu, Long-Kai Huang, Ying WeiNeurIPS 2022 · 18 citations
- Bayesian Meta-Learning for the Few-Shot Setting via Deep KernelsMassimiliano Patacchiola, Jack Turner, Elliot J. Crowley, Michael F. P. O'Boyle et al.NeurIPS 2020 · 167 citations
- Efficient and Effective Multi-task Grouping via Meta Learning on Task CombinationsXiaozhuang Song, Shun Zheng, Wei Cao, James J. Q. Yu et al.NeurIPS 2022 · 50 citations
- Meta Dropout: Learning to Perturb Latent Features for GeneralizationHaebeom Lee, Taewook Nam, Eunho Yang, Sung Ju HwangICLR 2020 · 59 citations
