Sharper Bounds for ℓp Sensitivity Sampling
David P. Woodruff, Taisuke Yasuda
摘要
In large scale machine learning, random sampling is a popular way to approximate datasets by a small representative subset of examples. In particular, sensitivity sampling is an intensely studied technique which provides provable guarantees on the quality of approximation, while reducing the number of examples to the product of the VC dimension d and the total sensitivity S in remarkably general settings. However, guarantees going beyond this general bound of Sd are known in perhaps only one setting, for ℓ2 subspace embeddings, despite intense study of sensitivity sampling in prior work. In this work, we show the first bounds for sensitivity sampling for ℓp subspace embeddings for p > 2 that improve over the general Sd bound, achieving a bound of roughly S 2-2/p for 2 < p < ∞. Furthermore, our techniques yield further new results in the study of sampling algorithms, showing that the root leverage score sampling algorithm achieves a bound of roughly d for 1 ≤ p < 2, and that a combination of leverage score and sensitivity sampling achieves an improved bound of roughly d 2/p S 2-4/p for 2 < p < ∞. Our sensitivity sampling results yield the best known sample complexity for a wide class of structured matrices that have small ℓp sensitivity. * Our sensitivity sampling results for p ∈ [1, 2) are subsumed by an earlier sharper result of Chen-Derezinski [CD21], which shows that the sensitivity scores upper bound Lewis weights up to a factor of d 1-p/2 , which implies a sampling complexity bound of Õ(ε -2 d 1-p/2 S). We refer to Section 1.1 as well as the recent work of [MO23] for more details. Question 1.4. How many samples are necessary for sensitivity sampling to output an ℓ p subspace embedding (2)? 3 We write Ax for the ith entry of the vector Ax ∈ R n . 4 We write Õ(f ) to denote f poly log f .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- On Generalization Bounds for Projective ClusteringMaria Sofia Bucarelli, Matilde Fjeldsø Larsen, Chris Schwiegelshohn, Mads ToftrupNeurIPS 2023 · 被引用 7 次
- Optimal bounds for ℓp sensitivity sampling via ℓ2 augmentationAlexander Munteanu, Simon OmlorICML 2024 · 被引用 6 次
- Computing Approximate 𝓁p SensitivitiesSwati Padmanabhan, David P. Woodruff, Richard ZhangNeurIPS 2023 · 被引用 5 次
- Turnstile ℓp leverage score sampling with applicationsAlexander Munteanu, Simon OmlorICML 2024 · 被引用 4 次
- Improving the Bit Complexity of Communication for Distributed Convex OptimizationMehrdad Ghadiri, Yin Tat Lee, Swati Padmanabhan, William Swartworth 等STOC 2024 · 被引用 1 次
它引用的顶会 Paper11
- Adversarial Robustness of Streaming Algorithms through Importance SamplingVladimir Braverman, Avinatan Hassidim, Yossi Matias, Mariano Schain 等NeurIPS 2021 · 被引用 56 次
- Coresets for Near-Convex FunctionsMurad Tukan, Alaa Maalouf, Dan FeldmanNeurIPS 2020 · 被引用 49 次
- Coresets for clustering in Euclidean spaces: importance sampling is nearly optimalLingxiao Huang, Nisheeth K. VishnoiSTOC 2020 · 被引用 36 次
- Near Optimal Linear Algebra in the Online and Sliding Window ModelsVladimir Braverman, Petros Drineas, Cameron Musco, Christopher Musco 等FOCS 2020 · 被引用 24 次
- Fast Regression for Structured InputsRaphael A. Meyer, Cameron Musco, Christopher Musco, David P. Woodruff 等ICLR 2022 · 被引用 14 次
相关 Paper
- Active Linear Regression for ℓp Norms and BeyondCameron Musco, Christopher Musco, David P. Woodruff, Taisuke YasudaFOCS 2022 · 被引用 4 次
- Root Ridge Leverage Score Sampling for ℓp Subspace ApproximationDavid P. Woodruff, Taisuke YasudaFOCS 2025 · 被引用 1 次
- Online Lewis Weight SamplingDavid P. Woodruff, Taisuke YasudaSODA 2023 · 被引用 3 次
- Robust Sparsification via SensitivityChansophea Wathanak In, Yi Li, David P. Woodruff, Xuan WuICML 2025
- Data-Efficient Learning via Clustering-Based Sensitivity Sampling: Foundation Models and BeyondKyriakos Axiotis, Vincent Cohen-Addad, Monika Henzinger, Sammy Jerome 等ICML 2024 · 被引用 19 次
