The continuous categorical: a novel simplex-valued exponential family
Elliott Gordon-Rodríguez, Gabriel Loaiza-Ganem, John P. Cunningham
Abstract
Simplex-valued data appear throughout statistics and machine learning, for example in the context of transfer learning and compression of deep networks. Existing models for this class of data rely on the Dirichlet distribution or other related loss functions; here we show these standard choices suffer systematically from a number of limitations, including bias and numerical issues that frustrate the use of flexible network models upstream of these distributions. We resolve these limitations by introducing a novel exponential family of distributions for modeling simplex-valued data -the continuous categorical, which arises as a nontrivial multivariate generalization of the recently discovered continuous Bernoulli. Unlike the Dirichlet and other typical choices, the continuous categorical results in a well-behaved probabilistic loss function that produces unbiased estimators, while preserving the mathematical simplicity of the Dirichlet. As well as exploring its theoretical properties, we introduce sampling methods for this distribution that are amenable to the reparameterization trick, and evaluate their performance. Lastly, we demonstrate that the continuous categorical outperforms standard choices empirically, across a simulation study, an applied example on multi-party elections, and a neural network compression task. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 020e866f-1708-4140-92cf-37d5bbb843ebCited by top-tier papers5
- Invertible Gaussian Reparameterization: Revisiting the Gumbel-SoftmaxAndres Potapczynski, Gabriel Loaiza-Ganem, John P. CunninghamNeurIPS 2020 · 44 citations
- Data Augmentation for Compositional Data: Advancing Predictive Models of the MicrobiomeElliott Gordon-Rodríguez, Thomas P. Quinn, John P. CunninghamNeurIPS 2022 · 15 citations
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 10 citations
- Learning Cut Generating Functions for Integer ProgrammingHongyu Cheng, Amitabh BasuNeurIPS 2024 · 10 citations
- Sparse Communication via Mixed DistributionsAntónio Farinhas, Wilker Aziz, Vlad Niculae, André F. T. MartinsICLR 2022 · 3 citations
Builds on1
Related papers
- Convolutional dictionary learning based auto-encoders for natural exponential-family distributionsBahareh Tolooshams, Andrew H. Song, Simona Temereanca, Demba E. BaICML 2020 · 26 citations
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae et al.NeurIPS 2020 · 55 citations
- Delving into Deep Imbalanced RegressionYuzhe Yang, Kaiwen Zha, Ying-Cong Chen, Hao Wang et al.ICML 2021 · 385 citations
- Bayesian Sparsification of Deep C-valued NetworksIvan Nazarov, Evgeny BurnaevICML 2020 · 4 citations
- Diffusing Gaussian Mixtures for Generating Categorical DataFlorence Regol, Mark CoatesAAAI 2023 · 6 citations
