Beyond Softmax: A Natural Parameterization for Categorical Random Variables
Alessandro Manenti, Cesare Alippi
摘要
Latent categorical variables are frequently found in deep learning architectures. They can model actions in discrete reinforcement-learning environments, represent categories in latent-variable models, or express relations in graph neural networks. Despite their widespread use, their discrete nature poses significant challenges to gradient-descent learning algorithms. While a substantial body of work has offered improved gradient estimation techniques, we take a complementary approach. Specifically, we: 1) revisit the ubiquitous softmax function and demonstrate its limitations from an information-geometric perspective; 2) replace the softmax with the catnat function, a function composed by a sequence of hierarchical binary splits; we prove that this choice offers significant advantages to gradient descent due to the resulting diagonal Fisher Information Matrix. A rich set of experiments - including graph structure learning, variational autoencoders, and reinforcement learning - empirically show that the proposed function improves the learning efficiency and yields models characterized by consistently higher test performance. Catnat is simple to implement and seamlessly integrates into existing codebases. Moreover, it remains compatible with standard training stabilization techniques and, as such, offers a better alternative to the softmax function.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- SLAPS: Self-Supervision Improves Structure Learning for Graph Neural NetworksBahare Fatemi, Layla El Asri, Seyed Mehran KazemiNeurIPS 2021 · 被引用 220 次
- Implicit MLE: Backpropagating Through Discrete Exponential Family DistributionsMathias Niepert, Pasquale Minervini, Luca FranceschiNeurIPS 2021 · 被引用 121 次
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
- Variational Inference for Graph Convolutional Networks in the Absence of Graph Data and Adversarial SettingsPantelis Elinas, Edwin V. Bonilla, Louis C. TiaoNeurIPS 2020 · 被引用 72 次
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 被引用 59 次
相关 Paper
- Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative TrickLennert De Smet, Emanuele Sansone, Pedro Zuidberg Dos MartiresNeurIPS 2023 · 被引用 17 次
- Categorical Reparameterization with Denoising Diffusion ModelsSamson Gourevitch, Alain Oliviero Durmus, Eric Moulines, Jimmy Olsson 等ICML 2026 · 被引用 1 次
- Coupled Gradient Estimators for Discrete Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2021 · 被引用 14 次
- Discrete Variational Autoencoding via Policy SearchMichael Drolet, Firas Al-Hafez, Aditya Bhatt, Jan Peters 等ICLR 2026
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 被引用 25 次
