Learning Group Importance using the Differentiable Hypergeometric Distribution
Thomas M. Sutter, Laura Manduchi, Alain Ryser, Julia E. Vogt
Abstract
Partitioning a set of elements into subsets of a priori unknown sizes is essential in many applications. These subset sizes are rarely explicitly learned - be it the cluster sizes in clustering applications or the number of shared versus independent generative latent factors in weakly-supervised learning. Probability distributions over correct combinations of subset sizes are non-differentiable due to hard constraints, which prohibit gradient-based optimization. In this work, we propose the differentiable hypergeometric distribution. The hypergeometric distribution models the probability of different group sizes based on their relative importance. We introduce reparameterizable gradients to learn the importance between groups and highlight the advantage of explicitly learning the size of subsets in two typical applications: weakly-supervised learning and clustering. In both applications, we outperform previous approaches, which rely on suboptimal heuristics to model the unknown size of groups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2eb07a62-74bc-4129-b96e-b6e6ec88b84fCited by top-tier papers3
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard et al.NeurIPS 2024 · 21 citations
- Differentiable Random Partition ModelsThomas M. Sutter, Alain Ryser, Joram Liebeskind, Julia E. VogtNeurIPS 2023 · 4 citations
- Estimating Unknown Population Sizes Using the Hypergeometric DistributionLiam Hodgson, Danilo BzdokICML 2024
Builds on3
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf et al.ICML 2020 · 361 citations
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
- Deep Conditional Gaussian Mixture Model for Constrained ClusteringLaura Manduchi, Kieran Chin-Cheong, Holger Michel, Sven Wellmann et al.NeurIPS 2021 · 40 citations
Related papers
- Categorical Reparameterization with Denoising Diffusion ModelsSamson Gourevitch, Alain Oliviero Durmus, Eric Moulines, Jimmy Olsson et al.ICML 2026 · 1 citation
- Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative TrickLennert De Smet, Emanuele Sansone, Pedro Zuidberg Dos MartiresNeurIPS 2023 · 17 citations
- SIMPLE: A Gradient Estimator for k-Subset SamplingKareem Ahmed, Zhe Zeng, Mathias Niepert, Guy Van den BroeckICLR 2023 · 2 citations
- Fast Generating A Large Number of Gumbel-Max VariablesYiyan Qi, Pinghui Wang, Yuanming Zhang, Junzhou Zhao et al.WWW 2020 · 6 citations
- Fiber Monte CarloNick Richardson, Deniz Oktay, Yaniv Ovadia, James C. Bowden et al.ICLR 2024
