Optimal approximate sampling from discrete probability distributions
Feras A. Saad, Cameron E. Freer, Martin C. Rinard, Vikash K. Mansinghka
Abstract
This paper addresses a fundamental problem in random variate generation: given access to a random source that emits a stream of independent fair bits, what is the most accurate and entropy-efficient algorithm for sampling from a discrete probability distribution (p 1 , . . . , p n ), where the probabilities of the output distribution ( p1 , . . . , pn ) of the sampling algorithm must be specified using at most k bits of precision? We present a theoretical framework for formulating this problem and provide new techniques for finding sampling algorithms that are optimal both statistically (in the sense of sampling accuracy) and information-theoretically (in the sense of entropy consumption). We leverage these results to build a system that, for a broad family of measures of statistical accuracy, delivers a sampling algorithm whose expected entropy usage is minimal among those that induce the same distribution (i.e., is "entropy-optimal") and whose output distribution ( p1 , . . . , pn ) is a closest approximation to the target distribution (p 1 , . . . , p n ) among all entropy-optimal sampling algorithms that operate within the specified k-bit precision. This optimal approximate sampler is also a closer approximation than any (possibly entropy-suboptimal) sampler that consumes a bounded amount of entropy with the specified precision, a class which includes floating-point implementations of inversion sampling and related methods found in many software libraries. We evaluate the accuracy, entropy consumption, precision requirements, and wall-clock runtime of our optimal approximate sampling algorithms on a broad set of distributions, demonstrating the ways that they are superior to existing approximate samplers and establishing that they often consume significantly fewer resources than are needed by exact samplers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d223409c-aa62-439e-a225-fa8e0d34d138Cited by top-tier papers4
- Fast Exact Leverage Score Sampling from Khatri-Rao Products with Applications to Tensor DecompositionVivek Bharadwaj, Osman Asif Malik, Riley Murray, Laura Grigori et al.NeurIPS 2023 · 14 citations
- Formally Verified Samplers from Probabilistic Programs with Loops and ConditioningAlexander Bagnall, Gordon Stewart, Anindya BanerjeePLDI 2023 · 5 citations
- Random Variate Generation with Formal GuaranteesFeras A. Saad, Wonyeol LeePLDI 2025
- Efficient Online Random Sampling via Randomness RecyclingThomas L. Draper, Feras A. SaadSODA 2026
Related papers
- Energy-Efficient Random Variate Generation via Compressed Lookup TablesJohann Ukrow, Anna Kazachkova, Nicolas Alder, Sven Köhler et al.ICLR 2026
- Instance-Optimality in I/O-Efficient Sampling and Sequential EstimationShyam Narayanan, Václav Rozhon, Jakub Tetek, Mikkel ThorupFOCS 2024
- TURF: Two-Factor, Universal, Robust, Fast Distribution Learning AlgorithmYi Hao, Ayush Jain, Alon Orlitsky, Vaishakh RavindrakumarICML 2022
- No Time to Hash: On Super-Efficient Entropy AccumulationYevgeniy Dodis, Siyao Guo, Noah Stephens-Davidowitz, Zhiye XieCRYPTO 2021 · 7 citations
- Learning Rate Free Bayesian Inference in Constrained DomainsLouis Sharrock, Lester Mackey, Christopher NemethNeurIPS 2023 · 3 citations
