Beyond Softmax: A Natural Parameterization for Categorical Random Variables
Alessandro Manenti, Cesare Alippi
Abstract
Latent categorical variables are frequently found in deep learning architectures. They can model actions in discrete reinforcement-learning environments, represent categories in latent-variable models, or express relations in graph neural networks. Despite their widespread use, their discrete nature poses significant challenges to gradient-descent learning algorithms. While a substantial body of work has offered improved gradient estimation techniques, we take a complementary approach. Specifically, we: 1) revisit the ubiquitous softmax function and demonstrate its limitations from an information-geometric perspective; 2) replace the softmax with the catnat function, a function composed by a sequence of hierarchical binary splits; we prove that this choice offers significant advantages to gradient descent due to the resulting diagonal Fisher Information Matrix. A rich set of experiments - including graph structure learning, variational autoencoders, and reinforcement learning - empirically show that the proposed function improves the learning efficiency and yields models characterized by consistently higher test performance. Catnat is simple to implement and seamlessly integrates into existing codebases. Moreover, it remains compatible with standard training stabilization techniques and, as such, offers a better alternative to the softmax function.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f80194b-9c3a-4204-83cb-0aa30ba142daBuilds on8
- SLAPS: Self-Supervision Improves Structure Learning for Graph Neural NetworksBahare Fatemi, Layla El Asri, Seyed Mehran KazemiNeurIPS 2021 · 220 citations
- Implicit MLE: Backpropagating Through Discrete Exponential Family DistributionsMathias Niepert, Pasquale Minervini, Luca FranceschiNeurIPS 2021 · 121 citations
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause et al.NeurIPS 2020 · 104 citations
- Variational Inference for Graph Convolutional Networks in the Absence of Graph Data and Adversarial SettingsPantelis Elinas, Edwin V. Bonilla, Louis C. TiaoNeurIPS 2020 · 72 citations
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
Related papers
- Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative TrickLennert De Smet, Emanuele Sansone, Pedro Zuidberg Dos MartiresNeurIPS 2023 · 17 citations
- Categorical Reparameterization with Denoising Diffusion ModelsSamson Gourevitch, Alain Oliviero Durmus, Eric Moulines, Jimmy Olsson et al.ICML 2026 · 1 citation
- Coupled Gradient Estimators for Discrete Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2021 · 14 citations
- Discrete Variational Autoencoding via Policy SearchMichael Drolet, Firas Al-Hafez, Aditya Bhatt, Jan Peters et al.ICLR 2026
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 25 citations
