Papers › Polynomial, trigonometric, and tropical activations

Polynomial, trigonometric, and tropical activations

3 Feb 2025arXiv:2502.01247archive 2025-07-28

Ismail Khalfaoui-Hassani, Stefan Kesselheim

Which functions can be used as activations in deep neural networks? This article explores families of functions based on orthonormal bases, including the Hermite polynomial basis and the Fourier trigonometric basis, as well as a basis resulting from the tropicalization of a polynomial basis. Our study shows that, through simple variance-preserving initialization and without additional clamping mechanisms, these activations can successfully be used to train deep models, such as GPT-2 for next-token prediction on OpenWebText and ConvNeXt for image classification on ImageNet. Our work addresses the issue of exploding and vanishing activations and gradients, particularly prevalent with polynomial activations, and opens the door for improving the efficiency of large-scale learning tasks. Furthermore, our approach provides insight into the structure of neural networks, revealing that networks with polynomial activations can be interpreted as multivariate polynomial mappings. Finally, using Hermite interpolation, we show that our activations can closely approximate classical ones in pre-trained models by matching both the function and its derivative, making them especially useful for fine-tuning tasks. These activations are available in the torchortho library, which can be accessed via: https://github.com/K-H-Ismail/torchortho.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

K-H-Ismail/torchortho officialmentioned in papermentioned on GitHubpytorchGPL-3.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image ClassificationLanguage ModellingText Generationimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification ImageNet ConvNeXt-T-Hermite Number of params 28M #547 of 1060 Archive leaderboard report
Image Classification ImageNet ConvNeXt-T-Hermite Top 1 Accuracy 82.34 #547 of 1060 Archive leaderboard report
Image Classification ImageNet ConvNeXt-T-Hermite Top 5 Accuracy 96.03 #547 of 1060 Archive leaderboard report
Language Modelling OpenWebText GPT2-Hermite eval_loss 2.91 #4 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-Hermite eval_perplexity 18.39 #4 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-Hermite parameters 124M #4 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-Tropical eval_loss 2.92 #5 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-Tropical eval_perplexity 18.64 #5 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-Tropical parameters 124M #5 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-Fourier eval_loss 2.93 #6 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-Fourier eval_perplexity 18.72 #6 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-Fourier parameters 124M #6 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-GELU eval_loss 2.95 #7 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-GELU eval_perplexity 19.24 #7 of 12 Archive leaderboard report
Language Modelling OpenWebText GPT2-GELU parameters 124M #7 of 12 Archive leaderboard report
Text Generation OpenWebText GPT2-Hermite eval_loss 2.91 #1 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBPEConvNeXtCosine AnnealingDense ConnectionsDiscriminative Fine-TuningDropoutGPT-2Hermite ActivationLayer NormalizationLinear LayerLinear Warmup With Cosine AnnealingMulti-Head AttentionResidual ConnectionSoftmaxWeight Decay

1 archive method tag without a method page not shown.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections