Supervised Kernel Thinning
Albert Gong, Kyuseong Choi, Raaz Dwivedi
摘要
The kernel thinning algorithm of Dwivedi&Mackey (2024) provides a better-than-i.i.d. compression of a generic set of points. By generating high-fidelity coresets of size significantly smaller than the input points, KT is known to speed up unsupervised tasks like Monte Carlo integration, uncertainty quantification, and non-parametric hypothesis testing, with minimal loss in statistical accuracy. In this work, we generalize the KT algorithm to speed up supervised learning problems involving kernel methods. Specifically, we combine two classical algorithms--Nadaraya-Watson (NW) regression or kernel smoothing, and kernel ridge regression (KRR)--with KT to provide a quadratic speed-up in both training and inference times. We show how distribution compression with KT in each setting reduces to constructing an appropriate kernel, and introduce the Kernel-Thinned NW and Kernel-Thinned KRR estimators. We prove that KT-based regression estimators enjoy significantly superior computational efficiency over the full-data estimators and improved statistical efficiency over i.i.d. subsampling of the training data. En route, we also provide a novel multiplicative error guarantee for compressing with KT. We validate our design choices with both simulations and real data experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- WildCat: Near-Linear Attention in Theory and PracticeTobias Schröder, Lester MackeyICML 2026 · 被引用 3 次
- Conditional Distribution Compression via the Kernel Conditional Mean EmbeddingDominic Broadbent, Nick Whiteley, Robert Allison, Tom LovettNeurIPS 2025 · 被引用 1 次
- Efficient and Accurate Explanation Estimation with Distribution CompressionHubert Baniecki, Giuseppe Casalicchio, Bernd Bischl, Przemyslaw BiecekICLR 2025
它引用的顶会 Paper2
相关 Paper
- Low-Rank ThinningAnnabelle Michael Carrell, Albert Gong, Abhishek Shetty, Raaz Dwivedi 等ICML 2025
- Debiased Distribution CompressionLingxiao Li, Raaz Dwivedi, Lester MackeyICML 2024 · 被引用 7 次
- Distributed Randomized Sketching Kernel LearningRong Yin, Yong Liu, Dan MengAAAI 2022 · 被引用 4 次
- Random Fourier Features via Fast Surrogate Leverage Weighted SamplingFanghui Liu, Xiaolin Huang, Yudong Chen, Jie Yang 等AAAI 2020 · 被引用 21 次
- Ridge Boosting is Both Robust and EfficientDavid Bruns-Smith, Zhongming Xie, Avi FellerNeurIPS 2025
