Does Preprocessing Help Training Over-parameterized Neural Networks?
Zhao Song, Shuo Yang, Ruizhe Zhang
摘要
Deep neural networks have achieved impressive performance in many areas. Designing a fast and provable method for training neural networks is a fundamental question in machine learning. The classical training method requires paying cost for both forward computation and backward computation, where is the width of the neural network, and we are given training points in -dimensional space. In this paper, we propose two novel preprocessing ideas to bypass this barrier: First, by preprocessing the initial weights of the neural networks, we can train the neural network in cost per iteration. Second, by preprocessing the input data points, we can train the neural network in cost per iteration. From the technical perspective, our result is a sophisticated combination of tools in different fields, greedy-type convergence analysis in optimization, sparsity observation in practical work, high-dimensional geometric search in data structure, concentration and anti-concentration in probability. Our results also provide theoretical insights for a large number of previously established fast training methods. In addition, our classical algorithm can be generalized to the Quantum computation model. Interestingly, we can get a similar sublinear cost per iteration but avoid preprocessing initial weights or input data points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- FL-NTK: A Neural Tangent Kernel-based Framework for Federated Learning AnalysisBaihe Huang, Xiaoxiao Li, Zhao Song, Xin YangICML 2021 · 被引用 66 次
- Federated Adversarial Learning: A Framework with Convergence AnalysisXiaoxiao Li, Zhao Song, Jiaming YangICML 2023 · 被引用 36 次
- Sketching for First Order Method: Efficient Algorithm for Low-Bandwidth Channel and VulnerabilityZhao Song, Yitan Wang, Zheng Yu, Lichen ZhangICML 2023 · 被引用 35 次
- Bypass Exponential Time Preprocessing: Fast Neural Network Training via Weight-Data Correlation PreprocessingJosh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang 等NeurIPS 2023 · 被引用 32 次
- Breaking the Linear Iteration Cost Barrier for Some Well-known Conditional Gradient Methods Using MaxIP Data-structuresZhaozhuo Xu, Zhao Song, Anshumali ShrivastavaNeurIPS 2021 · 被引用 32 次
它引用的顶会 Paper6
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- Quantum Algorithms for Deep Convolutional Neural NetworksIordanis Kerenidis, Jonas Landman, Anupam PrakashICLR 2020 · 被引用 163 次
- A Faster Interior Point Method for Semidefinite ProgrammingHaotian Jiang, Tarun Kathuria, Yin Tat Lee, Swati Padmanabhan 等FOCS 2020 · 被引用 62 次
- Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of DimensionalityYi Zhang, Orestis Plevrakis, Simon S. Du, Xingguo Li 等NeurIPS 2020 · 被引用 56 次
- An improved cutting plane method for convex optimization, convex-concave games, and its applicationsHaotian Jiang, Yin Tat Lee, Zhao Song, Sam Chiu-wai WongSTOC 2020 · 被引用 54 次
相关 Paper
- A Sublinear Adversarial Training AlgorithmYeqi Gao, Lianke Qin, Zhao Song, Yitan WangICLR 2024 · 被引用 27 次
- Modular Duality in Deep LearningJeremy Bernstein, Laker NewhouseICML 2025
- Deep Frequency Principle Towards Understanding Why Deeper Learning Is FasterZhiqin John Xu, Hanxu ZhouAAAI 2021 · 被引用 67 次
- Alternating Layered Variational Quantum Circuits Can Be Classically Optimized Efficiently Using Classical ShadowsAfrad Basheer, Yuan Feng, Christopher Ferrie, Sanjiang LiAAAI 2023 · 被引用 13 次
- The Effect of Weight Precision on the Neuron Count in Deep ReLU NetworksSonghua He, Periklis A. PapakonstantinouICML 2024
