Root Ridge Leverage Score Sampling for ℓp Subspace Approximation
David P. Woodruff, Taisuke Yasuda
摘要
The ℓp subspace approximation problem is an NP-hard low rank approximation problem that generalizes the median hyperplane problem (p = 1), principal component analysis (p = 2), and the center hyperplane problem (p = ∞). A popular approach to cope with the NP-hardness of this problem is to compute a strong coreset, which is a small weighted subset of the input points which simultaneously approximates the cost of every k-dimensional subspace, typically to (1 + ε) relative error for a small constant ε.
We obtain an algorithm for constructing a strong coreset for ℓp subspace approximation of size Õ(kε -4/p ) for p < 2 and Õ(k p/2 ε -p ) for p > 2. This offers the following improvements over prior work:
• We construct the first strong coresets with nearly optimal dependence on k for all p = 2. In prior work, [SW18] constructed coresets of modified points with a similar dependence on k, while [HV20] constructed true coresets with polynomially worse dependence on k.
• We recover or improve the best known ε dependence for all p. In particular, for p > 2, the [SW18] coreset of modified points had a dependence of ε -p 2 /2 and the [HV20] coreset had a dependence of ε -3p .
Our algorithm is based on sampling by root ridge leverage scores, which admits fast algorithms, especially for sparse or structured matrices. Our analysis completely avoids the use of the representative subspace theorem [SW18], which is a critical component of all prior dimension-independent coresets for ℓp subspace approximation.
Our techniques also lead to the first nearly optimal online strong coresets for ℓp subspace approximation with similar bounds as the offline setting, resolving a problem of [WY23a]. All prior approaches lose poly(k) factors in this setting, even when allowed to modify the original points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- Coresets for Classification - Simplified and StrengthenedTung Mai, Cameron Musco, Anup RaoNeurIPS 2021 · 被引用 39 次
- Coresets for clustering in Euclidean spaces: importance sampling is nearly optimalLingxiao Huang, Nisheeth K. VishnoiSTOC 2020 · 被引用 36 次
- Near Optimal Linear Algebra in the Online and Sliding Window ModelsVladimir Braverman, Petros Drineas, Cameron Musco, Christopher Musco 等FOCS 2020 · 被引用 24 次
- Coresets for Clustering in Excluded-minor Graphs and BeyondVladimir Braverman, Shaofeng H.-C. Jiang, Robert Krauthgamer, Xuan WuSODA 2021 · 被引用 21 次
- Towards optimal lower bounds for k-median and k-means coresetsVincent Cohen-Addad, Kasper Green Larsen, David Saulpic, Chris SchwiegelshohnSTOC 2022 · 被引用 20 次
相关 Paper
- High-Dimensional Geometric Streaming for Nearly Low Rank DataHossein Esfandiari, Praneeth Kacham, Vahab Mirrokni, David P. Woodruff 等ICML 2024 · 被引用 1 次
- Online Lewis Weight SamplingDavid P. Woodruff, Taisuke YasudaSODA 2023 · 被引用 3 次
- Coresets for Multiple ℓp RegressionDavid P. Woodruff, Taisuke YasudaICML 2024 · 被引用 3 次
- High-Dimensional Geometric Streaming in Polynomial SpaceDavid P. Woodruff, Taisuke YasudaFOCS 2022 · 被引用 3 次
- Tight Sensitivity Bounds For Smaller CoresetsAlaa Maalouf, Adiel Statman, Dan FeldmanKDD 2020 · 被引用 11 次
