Lune

FOCS2021Top-tier venue

On the Power of Preconditioning in Sparse Linear Regression

Jonathan A. Kelner, Frederic Koehler, Raghu Meka, Dhruv Rohatgi

2021Year
3Citations
6Top-tier citations

Abstract

Sparse linear regression is a fundamental problem in high-dimensional statistics, but strikingly little is known about how to efficiently solve it without restrictive conditions on the design matrix. We consider the (correlated) random design setting, where the covariates are independently drawn from a multivariate GaussianN(0, Σ)N(0,\ \Sigma), for somen×nn\times npositive semi-definite matrixΣ\Sigma, and seek estimatorsw^\hat{w}minimizing(w^−w∗)TΣ(w^−w∗)(\hat{w}-w^{\ast})^{T}\Sigma(\hat{w}-w^{\ast}), wherew∗w^{\ast}is the k-sparse ground truth. Information theoretically, one can achieve strong error bounds with onlyO(klog⁡n)O(k\log n)samples for arbitraryΣ\Sigmaandw∗w^{\ast}; however, no efficient algorithms are known to match these guarantees even witho(n)o(n)samples, without further assumptions onΣ\Sigmaorw∗w^{\ast}. Yet there is little evidence for this gap in the random design setting: computational lower bounds are only known for worst-case design matrices. To date, random-design instances (i.e. specific covariance matricesΣ\Sigma) have only been proven hard against the Lasso program and variants. More precisely, these “hard” instances can often be solved by Lasso after a simple change-of-basis (i.e. preconditioning). In this work, we give both upper and lower bounds clarifying the power of preconditioning as a tool for solving sparse linear regression problems. On the one hand, we show that the preconditioned Lasso can solve a large class of sparse linear regression problems nearly optimally: it succeeds whenever the dependency structure of the covariates, in the sense of the Markov property, has low treewidth - even ifΣ\Sigmais highly ill-conditioned. This upper bound builds on ideas from the wavelet and signal processing literature. As a special case of this result, we give an algorithm for sparse linear regression with covariates from an autoregressive time series model, where we also show that the (usual) Lasso provably fails. On the other hand, we construct (for the first time) random-design instances which are provably hard even for an optimally preconditioned Lasso. In fact, we complete our treewidth classification by proving that for any treewidth-t graph, there exists a Gaussian Markov Random Field on this graph such that the preconditioned Lasso, with any choice of preconditioner, requiresΩ(t1/20)\Omega(t^{1/20})samples to recoverO(log⁡n)O(\log n)-sparse signals when covariates are drawn from this model.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f06fd485-d509-4eb1-8ddf-46eb6d4ecf9c

Cited by top-tier papers6

Ask how each one uses it

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines