Multiple Hypothesis Dropout: Estimating the Parameters of Multi-Modal Output Distributions
David D. Nguyen, David Liebowitz, Salil S. Kanhere, Surya Nepal
Abstract
In many real-world applications, from robotics to pedestrian trajectory prediction, there is a need to predict multiple real-valued outputs to represent several potential scenarios. Current deep learning techniques to address multiple-output problems are based on two main methodologies: (1) mixture density networks, which suffer from poor stability at high dimensions, or (2) multiple choice learning (MCL), an approach that uses M single-output functions, each only producing a point estimate hypothesis. This paper presents a Mixture of Multiple-Output functions (MoM) approach using a novel variant of dropout, Multiple Hypothesis Dropout. Unlike traditional MCL-based approaches, each multiple-output function not only estimates the mean but also the variance for its hypothesis. This is achieved through a novel stochastic winner-take-all loss which allows each multiple-output function to estimate variance through the spread of its subnetwork predictions. Experiments on supervised learning problems illustrate that our approach outperforms existing solutions for reconstructing multimodal output distributions. Additional studies on unsupervised learning problems show that estimating the parameters of latent posterior distributions within a discrete autoencoder significantly improves codebook efficiency, sample quality, precision and recall.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Diverse Multimedia Layout Generation with Multi Choice LearningDavid D. Nguyen, Surya Nepal, Salil S. KanhereACM MM 2021 · 12 citations
- Taming Transformers for High-Resolution Image SynthesisPatrick Esser, Robin Rombach, Björn OmmerCVPR 2021
Related papers
- Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysisVictor Letzelter, Mathieu Fontaine, Mickaël Chen, Patrick Pérez et al.NeurIPS 2023 · 15 citations
- Bayesian Posterior Approximation With Stochastic EnsemblesOleksandr Balabanov, Bernhard Mehlig, Hampus LinanderCVPR 2023
- Estimating Epistemic and Aleatoric Uncertainty with a Single ModelMatthew Chan, Maria Molina, Chris MetzlerNeurIPS 2024 · 76 citations
- Uncertainty-Aware Audiovisual Activity Recognition Using Deep Bayesian Variational InferenceMahesh Subedar, Ranganath Krishnan, Paulo Lopez-Meyer, Omesh Tickoo et al.ICCV 2019 · 81 citations
- GFlowOut: Dropout with Generative Flow NetworksDianbo Liu, Moksh Jain, Bonaventure F. P. Dossou, Qianli Shen et al.ICML 2023 · 27 citations
