Approximation to Smooth Functions by Low-Rank Swish Networks
Zimeng Li, Hongjun Li, Jingyuan Wang, Ke Tang
摘要
While deep learning has witnessed remarkable achievements in a wide range of applications, its substantial computational cost imposes limitations on application scenarios of neural networks. To alleviate this problem, low-rank compression is proposed as a class of efficient and hardwarefriendly network compression methods, which reduce computation by replacing large matrices in neural networks with products of two small ones. In this paper, we implement low-rank networks by inserting a sufficiently narrow linear layer without bias between each of two adjacent nonlinear layers. We prove that low-rank Swish networks with a fixed depth are capable of approximating any function from the Hölder ball C β,R ([0, 1] d ) within an arbitrarily small error where β is the smooth parameter and R is the radius. Our proposed constructive approximation ensures that the width of linear hidden layers required for approximation is no more than one-third of the width of nonlinear layers, which implies that the computational cost can be decreased by at least one-third compared with a network with the same depth and width of nonlinear layers but without narrow linear hidden layers. Our theoretical finding can offer a theoretical basis for low-rank compression from the perspective of universal approximation theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language ModelsDayan Pan, Jingyuan Wang, Yilong Zhou, Jiawei Cheng 等AAAI 2026
- Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language ModelsDayan Pan, Jingyuan Wang, Xie YuKDD 2026
它引用的顶会 Paper14
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- PDFormer: Propagation Delay-Aware Dynamic Long-Range Transformer for Traffic Flow PredictionJiawei Jiang, Chengkai Han, Wayne Xin Zhao, Jingyuan WangAAAI 2023 · 被引用 542 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
相关 Paper
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise DecompositionLucas Liebenwein, Alaa Maalouf, Dan Feldman, Daniela RusNeurIPS 2021 · 被引用 60 次
- HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural NetworksJinqi Xiao, Chengming Zhang, Yu Gong, Miao Yin 等AAAI 2023 · 被引用 35 次
- Neural Network Approximation based on Hausdorff distance of Tropical ZonotopesPanagiotis Misiakos, Georgios Smyrnis, George Retsinas, Petros MaragosICLR 2022 · 被引用 10 次
- ALF: Autoencoder-based Low-rank Filter-sharing for Efficient Convolutional Neural NetworksAlexander Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Laura Hauenschild 等DAC 2020 · 被引用 5 次
- "Lossless" Compression of Deep Neural Networks: A High-dimensional Neural Tangent Kernel ApproachLingyu Gu, Yongqi Du, Yuan Zhang, Di Xie 等NeurIPS 2022 · 被引用 9 次
