Parallel Sampling via Counting
Nima Anari, Ruiquan Gao, Aviad Rubinstein
Abstract
We show how to use parallelization to speed up sampling from an arbitrary distribution µ on a product space [q]n, given oracle access to counting queries: ℙX∼ µ[XS=σS] for any S⊆ [n] and σS ∈ [q]S. Our algorithm takes O(n2/3· polylog(n,q)) parallel time, to the best of our knowledge, the first sublinear in n runtime for arbitrary distributions. Our results have implications for sampling in autoregressive models. Our algorithm directly works with an equivalent oracle that answers conditional marginal queries ℙX∼ µ[Xi=σi | XS=σS], whose role is played by a trained neural network in autoregressive models. This suggests a roughly n1/3-factor speedup is possible for sampling in any-order autoregressive models. We complement our positive result by showing a lower bound of Ω(n1/3) for the runtime of any parallel sampling algorithm making at most poly(n) queries to the counting oracle, even for q=2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2f2e1ce-4579-4495-8865-9d36926f990eCited by top-tier papers4
- Accelerating Diffusion LLMs via Adaptive Parallel DecodingDaniel Israel, Guy Van den Broeck, Aditya GroverNeurIPS 2025 · 114 citations
- Parallel Sampling via AutospeculationNima Anari, Carlo Baronio, CJ Chen, Alireza Haqi et al.STOC 2026 · 5 citations
- From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language ModelsHengyu Fu, Baihe Huang, Virginia Adams, Charles Wang et al.ICML 2026
- Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto SpeculationHengyuan Hu, Aniket Das, Dorsa Sadigh, Nima AnariICML 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Parallel Sampling of Diffusion ModelsAndy Shih, Suneel Belkhale, Stefano Ermon, Dorsa Sadigh et al.NeurIPS 2023 · 144 citations
- Training and Inference on Any-Order Autoregressive Models the Right WayAndy Shih, Dorsa Sadigh, Stefano ErmonNeurIPS 2022 · 68 citations
- Accelerating Feedforward Computation via Parallel Nonlinear Equation SolvingYang Song, Chenlin Meng, Renjie Liao, Stefano ErmonICML 2021 · 44 citations
Related papers
- Predictive Sampling with Forecasting Autoregressive ModelsAuke J. Wiggers, Emiel HoogeboomICML 2020 · 18 citations
- Predictive Querying for Autoregressive Neural Sequence ModelsAlex Boyd, Samuel Showalter, Stephan Mandt, Padhraic SmythNeurIPS 2022 · 6 citations
- Anytime Sampling for Autoregressive Models via Ordered AutoencodingYilun Xu, Yang Song, Sahaj Garg, Linyuan Gong et al.ICLR 2021 · 15 citations
- Parallel Simulation for Log-concave Sampling and Score-based Diffusion ModelsHuanjian Zhou, Masashi SugiyamaICML 2025
- Optimal Sublinear Sampling of Spanning Trees and Determinantal Point Processes via Average-Case Entropic IndependenceNima Anari, Yang P. Liu, Thuy-Duong VuongFOCS 2022 · 1 citation
