Lune

NeurIPS2025顶会

Online Learning of Neural Networks

Amit Daniely, Idan Mehalel, Elchanan Mossel

2025年份
9被引次数
1顶会引用

摘要

We study online learning of feedforward neural networks with the sign activation function that implement functions from the unit ball in Rd\mathbb{R}^d to a finite label set {1,…,Y}\{1, \ldots, Y\}. First, we characterize a margin condition that is sufficient and in some cases necessary for online learnability of a neural network: Every neuron in the first hidden layer classifies all instances with some margin γ\gamma bounded away from zero. Quantitatively, we prove that for any net, the optimal mistake bound is at most approximately TS(d,γ)\mathtt{TS}(d,\gamma), which is the (d,γ)(d,\gamma)-totally-separable-packing number, a more restricted variation of the standard (d,γ)(d,\gamma)-packing number. We complement this result by constructing a net on which any learner makes TS(d,γ)\mathtt{TS}(d,\gamma) many mistakes. We also give a quantitative lower bound of approximately TS(d,γ)≥max⁡{1/(γd)d,d}\mathtt{TS}(d,\gamma) \geq \max\{1/(\gamma \sqrt{d})^d, d\} when γ≥1/2\gamma \geq 1/2, implying that for some nets and input sequences every learner will err for exp⁡(d)\exp(d) many times, and that a dimension-free mistake bound is almost always impossible. To remedy this inevitable dependence on dd, it is natural to seek additional natural restrictions to be placed on the network, so that the dependence on dd is removed. We study two such restrictions. The first is the multi-index model, in which the function computed by the net depends only on k≪dk \ll d orthonormal directions. We prove a mistake bound of approximately (1.5/γ)k+2(1.5/\gamma)^{k + 2} in this model. The second is the extended margin assumption. In this setting, we assume that all neurons (in all layers) in the network classify every ingoing input from previous layer with margin γ\gamma bounded away from zero. In this model, we prove a mistake bound of approximately (log⁡Y)/γO(L)(\log Y)/ \gamma^{O(L)}, where L is the depth of the network.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖