Polynomial, trigonometric, and tropical activations
Ismail Khalfaoui Hassani, Stefan Kesselheim
Abstract
Which functions can be used as activations in deep neural networks? This article explores families of functions based on orthonormal bases, including the Hermite polynomial basis and the Fourier trigonometric basis, as well as a basis resulting from the tropicalization of a polynomial basis. Our study shows that, through simple variance-preserving initialization and without additional clamping mechanisms, these activations can successfully be used to train deep models, such as GPT-2 for next-token prediction on OpenWebText and ConvNeXt for image classification on ImageNet. Our work addresses the issue of exploding and vanishing activations and gradients, particularly prevalent with polynomial activations, and opens the door for improving the efficiency of large-scale learning tasks. Furthermore, our approach provides insight into the structure of neural networks, revealing that networks with polynomial activations can be interpreted as multivariate polynomial mappings. Finally, using Hermite interpolation, we show that our activations can closely approximate classical ones in pre-trained models by matching both the function and its derivative, making them especially useful for fine-tuning tasks. These activations are available in the torchortho 1 library. Recently, Yang & Wang (2025) employed the same principle to train learnable rational activations. However, they encountered a challenge: the second-order moment has no closed formulation in the case of rational fractions. The authors' solution for ensuring the convergence of such rational activation networks consisted in initializing them by fitting the polynomial coefficients to a classical 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72780949-3ff7-4841-a3c5-d6cab652ded5Builds on15
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep NetworksAlejandro Molina, Patrick Schramowski, Kristian KerstingICLR 2020 · 116 citations
- Git Re-Basin: Merging Models modulo Permutation SymmetriesSamuel K. Ainsworth, Jonathan Hayase, Siddhartha S. SrinivasaICLR 2023 · 32 citations
- Identifiability of Deep Polynomial Neural NetworksKonstantin Usevich, Ricardo Augusto Borsoi, Clara Dérand, Marianne ClauselNeurIPS 2025 · 21 citations
Related papers
- Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision ModelsJeffrey Wang, Jonathan Gregory, Grigorios ChrysosICML 2026
- Generating Accurate Pseudo-Labels in Semi-Supervised Learning and Avoiding Overconfident Predictions via Hermite Polynomial ActivationsVishnu Suresh Lokhande, Songwong Tasneeyapant, Abhay Venkatesh, Sathya N. Ravi et al.CVPR 2020
- Polynomial Composition Activations: Unleashing the Dynamics of Large Language ModelsZhijian Zhuo, Ya Wang, Yutao Zeng, Xiaoqing Li et al.ICLR 2025
- Regularization of polynomial networks for image recognitionGrigorios G. Chrysos, Bohan Wang, Jiankang Deng, Volkan CevherCVPR 2023
- Characterizing the spectrum of the NTK via a power series expansionMichael Murray, Hui Jin, Benjamin Bowman, Guido MontúfarICLR 2023 · 2 citations
