Optimal bounds for ℓp sensitivity sampling via ℓ2 augmentation
Alexander Munteanu, Simon Omlor
摘要
Data subsampling is one of the most natural methods to approximate a massively large data set by a small representative proxy. In particular, sensitivity sampling received a lot of attention, which samples points proportional to an individual importance measure called sensitivity. This framework reduces in very general settings the size of data to roughly the VC dimension times the total sensitivity while providing strong guarantees on the quality of approximation. The recent work of Woodruff&Yasuda (2023c) improved substantially over the general bound for the important problem of subspace embeddings to for . Their result was subsumed by an earlier bound which was implicitly given in the work of Chen&Derezinski (2021). We show that their result is tight when sampling according to plain sensitivities. We observe that by augmenting the sensitivities by sensitivities, we obtain better bounds improving over the aforementioned results to optimal linear sampling complexity for all . In particular, this resolves an open question of Woodruff&Yasuda (2023c) in the affirmative for and brings sensitivity subsampling into the regime that was previously only known to be possible using Lewis weights (Cohen&Peng, 2015). As an application of our main result, we also obtain an sensitivity sampling bound for logistic regression, where is a natural complexity measure for this problem. This improves over the previous bound of Mai et al. (2021) which was based on Lewis weights subsampling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Turnstile ℓp leverage score sampling with applicationsAlexander Munteanu, Simon OmlorICML 2024 · 被引用 4 次
- Data subsampling for Poisson regression with pth-root-linkHan Cheng Lie, Alexander MunteanuNeurIPS 2024 · 被引用 2 次
- The Space Complexity of Approximating Logistic LossGregory Dexter, Petros Drineas, Rajiv KhannaNeurIPS 2024 · 被引用 1 次
- On Coreset for LASSO Regression Problem with Sensitivity SamplingYuanbin Zou, Junyu Huang, Jianxin Wang, Qilong FengICLR 2026
它引用的顶会 Paper16
- Coresets for Near-Convex FunctionsMurad Tukan, Alaa Maalouf, Dan FeldmanNeurIPS 2020 · 被引用 49 次
- Improved Coresets for Euclidean k-MeansVincent Cohen-Addad, Kasper Green Larsen, David Saulpic, Chris Schwiegelshohn 等NeurIPS 2022 · 被引用 47 次
- Coresets for Classification - Simplified and StrengthenedTung Mai, Cameron Musco, Anup RaoNeurIPS 2021 · 被引用 39 次
- Oblivious Sketching for Logistic RegressionAlexander Munteanu, Simon Omlor, David P. WoodruffICML 2021 · 被引用 23 次
- Generic Coreset for Scalable Learning of Monotonic Kernels: Logistic Regression, Sigmoid and moreElad Tolochinsky, Ibrahim Jubran, Dan FeldmanICML 2022 · 被引用 19 次
相关 Paper
- Sharper Bounds for ℓp Sensitivity SamplingDavid P. Woodruff, Taisuke YasudaICML 2023 · 被引用 8 次
- Computing Approximate 𝓁p SensitivitiesSwati Padmanabhan, David P. Woodruff, Richard ZhangNeurIPS 2023 · 被引用 5 次
- Active Linear Regression for ℓp Norms and BeyondCameron Musco, Christopher Musco, David P. Woodruff, Taisuke YasudaFOCS 2022 · 被引用 4 次
- Online Lewis Weight SamplingDavid P. Woodruff, Taisuke YasudaSODA 2023 · 被引用 3 次
- Robust Sparsification via SensitivityChansophea Wathanak In, Yi Li, David P. Woodruff, Xuan WuICML 2025
