Root Ridge Leverage Score Sampling for ℓp Subspace Approximation
David P. Woodruff, Taisuke Yasuda
Abstract
The ℓp subspace approximation problem is an NP-hard low rank approximation problem that generalizes the median hyperplane problem (p = 1), principal component analysis (p = 2), and the center hyperplane problem (p = ∞). A popular approach to cope with the NP-hardness of this problem is to compute a strong coreset, which is a small weighted subset of the input points which simultaneously approximates the cost of every k-dimensional subspace, typically to (1 + ε) relative error for a small constant ε.
We obtain an algorithm for constructing a strong coreset for ℓp subspace approximation of size Õ(kε -4/p ) for p < 2 and Õ(k p/2 ε -p ) for p > 2. This offers the following improvements over prior work:
• We construct the first strong coresets with nearly optimal dependence on k for all p = 2. In prior work, [SW18] constructed coresets of modified points with a similar dependence on k, while [HV20] constructed true coresets with polynomially worse dependence on k.
• We recover or improve the best known ε dependence for all p. In particular, for p > 2, the [SW18] coreset of modified points had a dependence of ε -p 2 /2 and the [HV20] coreset had a dependence of ε -3p .
Our algorithm is based on sampling by root ridge leverage scores, which admits fast algorithms, especially for sparse or structured matrices. Our analysis completely avoids the use of the representative subspace theorem [SW18], which is a critical component of all prior dimension-independent coresets for ℓp subspace approximation.
Our techniques also lead to the first nearly optimal online strong coresets for ℓp subspace approximation with similar bounds as the offline setting, resolving a problem of [WY23a]. All prior approaches lose poly(k) factors in this setting, even when allowed to modify the original points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2fabdd0-1d6f-4de4-ba4f-009d404195ebCited by top-tier papers1
Ask how each one uses itBuilds on26
- Coresets for Classification - Simplified and StrengthenedTung Mai, Cameron Musco, Anup RaoNeurIPS 2021 · 39 citations
- Coresets for clustering in Euclidean spaces: importance sampling is nearly optimalLingxiao Huang, Nisheeth K. VishnoiSTOC 2020 · 36 citations
- Near Optimal Linear Algebra in the Online and Sliding Window ModelsVladimir Braverman, Petros Drineas, Cameron Musco, Christopher Musco et al.FOCS 2020 · 24 citations
- Coresets for Clustering in Excluded-minor Graphs and BeyondVladimir Braverman, Shaofeng H.-C. Jiang, Robert Krauthgamer, Xuan WuSODA 2021 · 21 citations
- Towards optimal lower bounds for k-median and k-means coresetsVincent Cohen-Addad, Kasper Green Larsen, David Saulpic, Chris SchwiegelshohnSTOC 2022 · 20 citations
Related papers
- High-Dimensional Geometric Streaming for Nearly Low Rank DataHossein Esfandiari, Praneeth Kacham, Vahab Mirrokni, David P. Woodruff et al.ICML 2024 · 1 citation
- Online Lewis Weight SamplingDavid P. Woodruff, Taisuke YasudaSODA 2023 · 3 citations
- Coresets for Multiple ℓp RegressionDavid P. Woodruff, Taisuke YasudaICML 2024 · 3 citations
- High-Dimensional Geometric Streaming in Polynomial SpaceDavid P. Woodruff, Taisuke YasudaFOCS 2022 · 3 citations
- Tight Sensitivity Bounds For Smaller CoresetsAlaa Maalouf, Adiel Statman, Dan FeldmanKDD 2020 · 11 citations
