Evidential Softmax for Sparse Multimodal Distributions in Deep Generative Models
Phil Chen, Masha Itkina, Ransalu Senanayake, Mykel J. Kochenderfer
Abstract
Many applications of generative models rely on the marginalization of their highdimensional output probability distributions. Normalization functions that yield sparse probability distributions can make exact marginalization more computationally tractable. However, sparse normalization functions usually require alternative loss functions for training since the log-likelihood is undefined for sparse probability distributions. Furthermore, many sparse normalization functions often collapse the multimodality of distributions. In this work, we present ev-softmax, a sparse normalization function that preserves the multimodality of probability distributions. We derive its properties, including its gradient in closed-form, and introduce a continuous family of approximations to ev-softmax that have full support and can be trained with probabilistic loss functions such as negative log-likelihood and Kullback-Leibler divergence. We evaluate our method on a variety of generative models, including variational autoencoders and auto-regressive architectures. Our method outperforms existing dense and sparse normalization techniques in distributional accuracy. We demonstrate that ev-softmax successfully reduces the dimensionality of probability distributions while maintaining multimodality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07d5ec13-6243-4b75-841b-6964b72eab6bCited by top-tier papers5
- Adaptive Compositional Continual Meta-LearningBin Wu, Jinyuan Fang, Xiangxiang Zeng, Shangsong Liang et al.ICML 2023 · 12 citations
- Failures Are Fated, But Can Be Faded: Characterizing and Mitigating Unwanted Behaviors in Large-Scale Vision and Language ModelsSom Sagar, Aditya Taparia, Ransalu SenanayakeICML 2024 · 11 citations
- Learning Discrete Structured Variational Auto-Encoder using Natural Evolution StrategiesAlon Berliner, Guy Rotman, Yossi Adi, Roi Reichart et al.ICLR 2022 · 5 citations
- MultiMax: Sparse and Multi-Modal Attention LearningYuxuan Zhou, Mario Fritz, Margret KeuperICML 2024 · 4 citations
- Random-Set Neural NetworksShireen Kudukkil Manchingal, Muhammad Mubashar, Kaizheng Wang, Keivan Shariatmadar et al.ICLR 2025
Builds on2
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 25 citations
- Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational AutoencodersMasha Itkina, Boris Ivanovic, Ransalu Senanayake, Mykel J. Kochenderfer et al.NeurIPS 2020 · 21 citations
Related papers
- SurVAE Flows: Surjections to Bridge the Gap between VAEs and FlowsDidrik Nielsen, Priyank Jaini, Emiel Hoogeboom, Ole Winther et al.NeurIPS 2020 · 100 citations
- A Batch Normalized Inference Network Keeps the KL Vanishing AwayQile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma et al.ACL 2020 · 70 citations
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae et al.NeurIPS 2020 · 55 citations
- Generative Marginalization ModelsSulin Liu, Peter J. Ramadge, Ryan P. AdamsICML 2024 · 3 citations
- Improving Variational Autoencoders with Density Gap-based RegularizationJianfei Zhang, Jun Bai, Chenghua Lin, Yanmeng Wang et al.NeurIPS 2022 · 11 citations
