Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep Networks
Alejandro Molina, Patrick Schramowski, Kristian Kersting
Abstract
The performance of deep network learning strongly depends on the choice of the non-linear activation function associated with each neuron. However, deciding on the best activation is non-trivial, and the choice depends on the architecture, hyper-parameters, and even on the dataset. Typically these activations are fixed by hand before training. Here, we demonstrate how to eliminate the reliance on first picking fixed activation functions by using flexible parametric rational functions instead. The resulting Pade Activation Units (PAUs) can both approximate common activation functions and also learn new ones while providing compact representations. Our empirical evidence shows that end-to-end learning deep networks with PAUs can increase the predictive performance. Moreover, PAUs pave the way to approximations with provable robustness. this https URL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 996f8a5d-ed58-4621-be52-1156b7a0ca59Cited by top-tier papers18
- Rational neural networksNicolas Boullé, Yuji Nakatsukasa, Alex TownsendNeurIPS 2020 · 130 citations
- Adaptive Rational Activations to Boost Deep Reinforcement LearningQuentin Delfosse, Patrick Schramowski, Martin Mundt, Alejandro Molina et al.ICLR 2024 · 25 citations
- Learning specialized activation functions with the Piecewise Linear UnitYucong Zhou, Zezhou Zhu, Zhao ZhongICCV 2021 · 17 citations
- ErfAct and Pserf: Non-monotonic Smooth Trainable Activation FunctionsKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyAAAI 2022 · 16 citations
- CoFrNets: Interpretable Neural Architecture Inspired by Continued FractionsIsha Puri, Amit Dhurandhar, Tejaswini Pedapati, Karthikeyan Shanmugam et al.NeurIPS 2021 · 13 citations
Related papers
- Differential Equation Units: Learning Functional Forms of Activation Functions from DataMohamadAli Torkamani, Shiv Shankar, Amirmohammad Rooshenas, Phillip WallisAAAI 2020 · 3 citations
- Rational Neural Networks have Expressivity AdvantagesMaosen Tang, Alex TownsendICML 2026 · 1 citation
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 10 citations
- Deep Learning for Functional Data Analysis with Adaptive Basis LayersJunwen Yao, Jonas Mueller, Jane-Ling WangICML 2021 · 40 citations
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
