Lune

NeurIPS2020顶会

Debiasing Distributed Second Order Optimization with Surrogate Sketching and Scaled Regularization

Michal Derezinski, Burak Bartan, Mert Pilanci, Michael W. Mahoney

2020年份
28被引次数
7顶会引用

摘要

In distributed second order optimization, a standard strategy is to average many local estimates, each of which is based on a small sketch or batch of the data. However, the local estimates on each machine are typically biased, relative to the full solution on all of the data, and this can limit the effectiveness of averaging. Here, we introduce a new technique for debiasing the local estimates, which leads to both theoretical and empirical improvements in the convergence rate of distributed second order methods. Our technique has two novel components: (1) modifying standard sketching techniques to obtain what we call a surrogate sketch; and (2) carefully scaling the global regularization parameter for local computations. Our surrogate sketches are based on determinantal point processes, a family of distributions for which the bias of an estimate of the inverse Hessian can be computed exactly. Based on this computation, we show that when the objective being minimized is l2l_2-regularized with parameter λ\lambda and individual machines are each given a sketch of size mm, then to eliminate the bias, local estimates should be computed using a shrunk regularization parameter given by λ′=λ⋅(1−dλm)\lambda^{\prime}=\lambda\cdot(1-\frac{d_{\lambda}}{m}), where dλd_{\lambda} is the λ\lambda-effective dimension of the Hessian (or, for quadratic problems, the data matrix).

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 4fe0fad9-c7e1-4078-a3cd-d4ea4a313fc6

引用它的顶会 Paper7

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖