Rational Neural Networks have Expressivity Advantages
Maosen Tang, Alex Townsend
摘要
We study neural networks with trainable low-degree rational activation functions and show that they are more expressive and parameter-efficient than modern piecewise-linear and smooth activations such as ELU, LeakyReLU, LogSigmoid, PReLU, ReLU, SELU, CELU, Sigmoid, SiLU, Mish, Softplus, Tanh, Softmin, Softmax, and LogSoftmax. For an error target of , we establish approximation-theoretic separations: Any network built from standard fixed activations can be uniformly approximated on compact domains by a rational-activation network with only overhead in size, while the converse provably requires parameters in the worst case. This exponential gap persists at the level of full networks and extends to gated activations and transformer-style nonlinearities. In practice, rational activations integrate seamlessly into standard architectures and training pipelines, allowing rationals to match or outperform fixed activations under identical architectures and optimizers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
相关 Paper
- Rational neural networksNicolas Boullé, Yuji Nakatsukasa, Alex TownsendNeurIPS 2020 · 被引用 130 次
- Floating-Point Neural Networks Can Represent Almost All Floating-Point FunctionsGeonho Hwang, Yeachan Park, Wonyeol Lee, Sejun ParkICML 2025
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep NetworksAlejandro Molina, Patrick Schramowski, Kristian KerstingICLR 2020 · 被引用 116 次
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 被引用 10 次
- Minimum width for universal approximation using ReLU networks on compact domainNamjun Kim, Chanho Min, Sejun ParkICLR 2024 · 被引用 19 次
