Beyond Worst-Case Dimensionality Reduction for Sparse Vectors
Sandeep Silwal, David P. Woodruff, Qiuyi Zhang
Abstract
We study beyond worst-case dimensionality reduction for s-sparse vectors (vectors with at most s non-zero coordinates). Our work is divided into two parts, each focusing on a different facet of beyond worst-case analysis: (a) We first consider average-case guarantees for embedding s-sparse vectors. Here, a well-known folklore upper bound based on the birthday-paradox states: For any collection X of s-sparse vectors in R d , there exists a linear map A : R d → R O(s 2 ) which exactly preserves the norm of 99% of the vectors in X in any ℓ p norm (as opposed to the usual setting where guarantees hold for all vectors). We provide novel lower bounds showing that this is indeed optimal in many settings. Specifically, any oblivious linear map satisfying similar average-case guarantees must map to Ω(s 2 ) dimensions. The same lower bound also holds for a wider class of sufficiently smooth maps, including 'encoder-decoder schemes', where we compare the norm of the original vector to that of a smooth function of the embedding. These lower bounds reveal a surprising separation result for smooth embeddings of sparse vectors, as an upper bound of O(s log(d)) is possible if we instead use arbitrary functions, e.g., via compressed sensing algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fd6cad5-ab06-428a-8a09-b6d0501ec891Builds on7
- Dimensionality Reduction for Wasserstein BarycenterZachary Izzo, Sandeep Silwal, Samson ZhouNeurIPS 2021 · 25 citations
- Randomized Dimensionality Reduction for Facility Location and Single-Linkage ClusteringShyam Narayanan, Sandeep Silwal, Piotr Indyk, Or ZamirICML 2021 · 16 citations
- Hardness and Algorithms for Robust and Sparse OptimizationEric Price, Sandeep Silwal, Samson ZhouICML 2022 · 10 citations
- Terminal Embeddings in Sublinear TimeYeshwanth Cherapanamjeri, Jelani NelsonFOCS 2021 · 5 citations
- Streaming Euclidean Max-Cut: Dimension vs Data ReductionXiaoyu Chen, Shaofeng H.-C. Jiang, Robert KrauthgamerSTOC 2023 · 4 citations
Related papers
- Sparse Dimensionality Reduction RevisitedMikael Møller Høgsgaard, Lior Kamma, Kasper Green Larsen, Jelani Nelson et al.ICML 2024 · 3 citations
- Optimal Embedding Dimension for Sparse Subspace EmbeddingsShabarish Chenakkod, Michal Derezinski, Xiaoyu Dong, Mark RudelsonSTOC 2024 · 5 citations
- Tight Bounds for the Subspace Sketch Problem with ApplicationsYi Li, Ruosong Wang, David P. WoodruffSODA 2020 · 5 citations
- How to Compress Encrypted DataNils Fleischhacker, Kasper Green Larsen, Mark SimkinEUROCRYPT 2023 · 5 citations
- Robust Testing in High-Dimensional Sparse ModelsAnand Jerry George, Clément L. CanonneNeurIPS 2022 · 4 citations
