Sampling Multimodal Distributions with the Vanilla Score: Benefits of Data-Based Initialization
Frederic Koehler, Thuy-Duong Vuong
Abstract
There is a long history, as well as a recent explosion of interest, in statistical and generative modeling approaches based on score functions -- derivatives of the log-likelihood of a distribution. In seminal works, Hyvärinen proposed vanilla score matching as a way to learn distributions from data by computing an estimate of the score function of the underlying ground truth, and established connections between this method and established techniques like Contrastive Divergence and Pseudolikelihood estimation. It is by now well-known that vanilla score matching has significant difficulties learning multimodal distributions. Although there are various ways to overcome this difficulty, the following question has remained unanswered -- is there a natural way to sample multimodal distributions using just the vanilla score? Inspired by a long line of related experimental works, we prove that the Langevin diffusion with early stopping, initialized at the empirical distribution, and run on a score function estimated from data successfully generates natural multimodal distributions (mixtures of log-concave distributions).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f7d4673-9c1d-4c11-be02-75c85abeed14Cited by top-tier papers6
- Critical windows: non-asymptotic theory for feature emergence in diffusion modelsMarvin Li, Sitan ChenICML 2024 · 34 citations
- FOOGD: Federated Collaboration for Both Out-of-distribution Generalization and DetectionXinting Liao, Weiming Liu, Pengyang Zhou, Fengyuan Yu et al.NeurIPS 2024 · 24 citations
- Weak Poincaré Inequalities, Simulated Annealing, and Sampling from Spherical Spin GlassesBrice Huang, Sidhanth Mohanty, Amit Rajaraman, David X. WuSTOC 2025 · 13 citations
- Computational Bottlenecks for Denoising DiffusionsViet Vu, Andrea MontanariICLR 2026 · 3 citations
- On the Robustness of Langevin Dynamics to Score Function ErrorDaniel Cao, August Chen, Karthik Sridharan, Yuchen WuICML 2026 · 2 citations
Builds on5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Convergence for score-based generative modeling with polynomial complexityHolden Lee, Jianfeng Lu, Yixin TanNeurIPS 2022 · 221 citations
- On the Anatomy of MCMC-Based Maximum Likelihood Learning of Energy-Based ModelsErik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu et al.AAAI 2020 · 182 citations
- Statistical Efficiency of Score Matching: The View from IsoperimetryFrederic Koehler, Alexander Heckett, Andrej RisteskiICLR 2023 · 6 citations
Related papers
- Particle Denoising Diffusion SamplerAngus Phillips, Hai-Dang Dau, Michael John Hutchinson, Valentin De Bortoli et al.ICML 2024 · 60 citations
- Score-Based Generative Modeling with Critically-Damped Langevin DiffusionTim Dockhorn, Arash Vahdat, Karsten KreisICLR 2022 · 276 citations
- Score-Based Diffusion meets Annealed Importance SamplingArnaud Doucet, Will Grathwohl, Alexander G. de G. Matthews, Heiko StrathmannNeurIPS 2022 · 68 citations
- Provable Convergence and Limitations of Geometric Tempering for Langevin DynamicsOmar Chehab, Anna Korba, Austin J. Stromme, Adrien VacherICLR 2025
- Adversarial score matching and improved sampling for image generationAlexia Jolicoeur-Martineau, Rémi Piché-Taillefer, Ioannis Mitliagkas, Remi Tachet des CombesICLR 2021 · 137 citations
