Deep Network Approximation in Terms of Intrinsic Parameters
Zuowei Shen, Haizhao Yang, Shijun Zhang
摘要
One of the arguments to explain the success of deep learning is the powerful approximation capacity of deep neural networks. Such capacity is generally accompanied by the explosive growth of the number of parameters, which, in turn, leads to high computational costs. It is of great interest to ask whether we can achieve successful deep learning with a small number of learnable parameters adapting to the target function. From an approximation perspective, this paper shows that the number of parameters that need to be learned can be significantly smaller than people typically expect. First, we theoretically design ReLU networks with a few learnable parameters to achieve an attractive approximation. We prove by construction that, for any Lipschitz continuous function on with a Lipschitz constant , a ReLU network with intrinsic parameters (those depending on ) can approximate with an exponentially small error . Such a result is generalized to generic continuous functions. Furthermore, we show that the idea of learning a small number of parameters to achieve a good approximation can be numerically observed. We conduct several experiments to verify that training a small part of parameters can also achieve good results for classification problems if other parameters are pre-specified or pre-trained from a related problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Neural Network Architecture Beyond Width and DepthShijun Zhang, Zuowei Shen, Haizhao YangNeurIPS 2022 · 被引用 25 次
- On Enhancing Expressive Power via Compositions of Single Fixed-Size ReLU NetworkShijun Zhang, Jianfeng Lu, Hongkai ZhaoICML 2023 · 被引用 9 次
它引用的顶会 Paper4
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- ThunderNet: Towards Real-Time Generic Object Detection on Mobile DevicesZheng Qin, Zeming Li, Zhaoning Zhang, Yiping Bao 等ICCV 2019 · 被引用 282 次
- Elementary superexpressive activationsDmitry YarotskyICML 2021 · 被引用 46 次
相关 Paper
- ReLU Network with Width d+O(1) Can Achieve Optimal Approximation RateChenghao Liu, Minghua ChenICML 2024 · 被引用 3 次
- Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During TrainingMax Milkert, David Hyde, Forrest J. LaineICML 2025
- Most Activation Functions Can Win the Lottery Without Excessive DepthRebekka BurkholzNeurIPS 2022 · 被引用 27 次
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 被引用 8 次
- Shallow and Deep Networks are Near-Optimal Approximators of Korobov FunctionsMoïse Blanchard, Mohammed Amine BennounaICLR 2022 · 被引用 11 次
