Estimating Unknown Population Sizes Using the Hypergeometric Distribution
Liam Hodgson, Danilo Bzdok
摘要
The multivariate hypergeometric distribution describes sampling without replacement from a discrete population of elements divided into multiple categories. Addressing a gap in the literature, we tackle the challenge of estimating discrete distributions when both the total population size and the sizes of its constituent categories are unknown. Here, we propose a novel solution using the hypergeometric likelihood to solve this estimation challenge, even in the presence of severe under-sampling. We develop our approach to account for a data generating process where the ground-truth is a mixture of distributions conditional on a continuous latent variable, such as with collaborative filtering, using the variational autoencoder framework. Empirical data simulation demonstrates that our method outperforms other likelihood functions used to model count data, both in terms of accuracy of population size estimate and in its ability to learn an informative latent space. We demonstrate our method's versatility through applications in NLP, by inferring and estimating the complexity of latent vocabularies in text excerpts, and in biology, by accurately recovering the true number of gene transcripts from sparse single-cell genomics data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Confidence sequences for sampling without replacementIan Waudby-Smith, Aaditya RamdasNeurIPS 2020 · 被引用 57 次
- On the benefits of maximum likelihood estimation for Regression and ForecastingPranjal Awasthi, Abhimanyu Das, Rajat Sen, Ananda Theertha SureshICLR 2022 · 被引用 14 次
- Learning Group Importance using the Differentiable Hypergeometric DistributionThomas M. Sutter, Laura Manduchi, Alain Ryser, Julia E. VogtICLR 2023 · 被引用 1 次
相关 Paper
- Cooperation in the Latent Space: The Benefits of Adding Mixture Components in Variational AutoencodersOskar Kviman, Ricky Molén, Alexandra Hotti, Semih Kurt 等ICML 2023 · 被引用 16 次
- Negative Binomial Variational Autoencoders for Overdispersed Latent ModelingYixuan Zhang, Jinhao Sheng, Wenxin Zhang, Quyu Kong 等CVPR 2026 · 被引用 3 次
- Poisson Variational AutoencoderHadi Vafaii, Dekel Galor, Jacob L. YatesNeurIPS 2024 · 被引用 18 次
- Decision-Making with Auto-Encoding Variational BayesRomain Lopez, Pierre Boyeau, Nir Yosef, Michael I. Jordan 等NeurIPS 2020 · 被引用 22,845 次
- Improving black-box optimization in VAE latent space using decoder uncertaintyPascal Notin, José Miguel Hernández-Lobato, Yarin GalNeurIPS 2021 · 被引用 76 次
