Lune

NeurIPS2021顶会

Does Preprocessing Help Training Over-parameterized Neural Networks?

Zhao Song, Shuo Yang, Ruizhe Zhang

2021年份
52被引次数
14顶会引用

摘要

Deep neural networks have achieved impressive performance in many areas. Designing a fast and provable method for training neural networks is a fundamental question in machine learning. The classical training method requires paying Ω(mnd)\Omega(mnd) cost for both forward computation and backward computation, where mm is the width of the neural network, and we are given nn training points in dd-dimensional space. In this paper, we propose two novel preprocessing ideas to bypass this Ω(mnd)\Omega(mnd) barrier: ∙\bullet First, by preprocessing the initial weights of the neural networks, we can train the neural network in O~(m1−Θ(1/d)nd)\widetilde{O}(m^{1-\Theta(1/d)} n d) cost per iteration. ∙\bullet Second, by preprocessing the input data points, we can train the neural network in O~(m4/5nd)\widetilde{O} (m^{4/5} nd ) cost per iteration. From the technical perspective, our result is a sophisticated combination of tools in different fields, greedy-type convergence analysis in optimization, sparsity observation in practical work, high-dimensional geometric search in data structure, concentration and anti-concentration in probability. Our results also provide theoretical insights for a large number of previously established fast training methods. In addition, our classical algorithm can be generalized to the Quantum computation model. Interestingly, we can get a similar sublinear cost per iteration but avoid preprocessing initial weights or input data points.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext ebf4d6f1-2063-4f0f-86b9-1d40c41c9324

引用它的顶会 Paper14

问问它们各自怎么用它

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖