The continuous categorical: a novel simplex-valued exponential family
Elliott Gordon-Rodríguez, Gabriel Loaiza-Ganem, John P. Cunningham
摘要
Simplex-valued data appear throughout statistics and machine learning, for example in the context of transfer learning and compression of deep networks. Existing models for this class of data rely on the Dirichlet distribution or other related loss functions; here we show these standard choices suffer systematically from a number of limitations, including bias and numerical issues that frustrate the use of flexible network models upstream of these distributions. We resolve these limitations by introducing a novel exponential family of distributions for modeling simplex-valued data -the continuous categorical, which arises as a nontrivial multivariate generalization of the recently discovered continuous Bernoulli. Unlike the Dirichlet and other typical choices, the continuous categorical results in a well-behaved probabilistic loss function that produces unbiased estimators, while preserving the mathematical simplicity of the Dirichlet. As well as exploring its theoretical properties, we introduce sampling methods for this distribution that are amenable to the reparameterization trick, and evaluate their performance. Lastly, we demonstrate that the continuous categorical outperforms standard choices empirically, across a simulation study, an applied example on multi-party elections, and a neural network compression task. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Invertible Gaussian Reparameterization: Revisiting the Gumbel-SoftmaxAndres Potapczynski, Gabriel Loaiza-Ganem, John P. CunninghamNeurIPS 2020 · 被引用 44 次
- Data Augmentation for Compositional Data: Advancing Predictive Models of the MicrobiomeElliott Gordon-Rodríguez, Thomas P. Quinn, John P. CunninghamNeurIPS 2022 · 被引用 15 次
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 被引用 10 次
- Learning Cut Generating Functions for Integer ProgrammingHongyu Cheng, Amitabh BasuNeurIPS 2024 · 被引用 10 次
- Sparse Communication via Mixed DistributionsAntónio Farinhas, Wilker Aziz, Vlad Niculae, André F. T. MartinsICLR 2022 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Convolutional dictionary learning based auto-encoders for natural exponential-family distributionsBahareh Tolooshams, Andrew H. Song, Simona Temereanca, Demba E. BaICML 2020 · 被引用 26 次
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae 等NeurIPS 2020 · 被引用 55 次
- Delving into Deep Imbalanced RegressionYuzhe Yang, Kaiwen Zha, Ying-Cong Chen, Hao Wang 等ICML 2021 · 被引用 385 次
- Bayesian Sparsification of Deep C-valued NetworksIvan Nazarov, Evgeny BurnaevICML 2020 · 被引用 4 次
- Diffusing Gaussian Mixtures for Generating Categorical DataFlorence Regol, Mark CoatesAAAI 2023 · 被引用 6 次
