On Divergence Measures for Bayesian Pseudocoresets
Balhae Kim, Jungwon Choi, Seanie Lee, Yoonho Lee, Jung-Woo Ha, Juho Lee
Abstract
A Bayesian pseudocoreset is a small synthetic dataset for which the posterior over parameters approximates that of the original dataset. While promising, the scalability of Bayesian pseudocoresets is not yet validated in realistic problems such as image classification with deep neural networks. On the other hand, dataset distillation methods similarly construct a small dataset such that the optimization using the synthetic dataset converges to a solution with performance competitive with optimization using full data. Although dataset distillation has been empirically verified in large-scale settings, the framework is restricted to point estimates, and their adaptation to Bayesian inference has not been explored. This paper casts two representative dataset distillation algorithms as approximations to methods for constructing pseudocoresets by minimizing specific divergence measures: reverse KL divergence and Wasserstein distance. Furthermore, we provide a unifying view of such divergence measures in Bayesian pseudocoreset construction. Finally, we propose a novel Bayesian pseudocoreset algorithm based on minimizing forward KL divergence. Our empirical results demonstrate that the pseudocoresets constructed from these methods reflect the true posterior even in high-dimensional Bayesian inference problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7084f87b-7091-4fc4-9929-90784e152935Cited by top-tier papers5
- Multisize Dataset CondensationYang He, Lingao Xiao, Joey Tianyi Zhou, Ivor W. TsangICLR 2024 · 22 citations
- Low-Rank Similarity Mining for Multimodal Dataset DistillationYue Xu, Zhilin Lin, Yusong Qiu, Cewu Lu et al.ICML 2024 · 14 citations
- Function Space Bayesian Pseudocoreset for Bayesian Neural NetworksBalhae Kim, Hyungi Lee, Juho LeeNeurIPS 2023 · 3 citations
- CHESS: Chebyshev Spectral Synthesis for Trajectory CondensationRuituo Wu, Hongyu Zhang, Qiang Wang, Jiawei Du et al.ICML 2026
- Variational Bayesian Pseudo-CoresetHyungi Lee, Seungyoo Lee, Juho LeeICLR 2025
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
Related papers
- Bayesian PseudocoresetsDionysis Manousakas, Zuheng Xu, Cecilia Mascolo, Trevor CampbellNeurIPS 2020 · 35 citations
- Fast Bayesian Coresets via Subsampling and Quasi-Newton RefinementCian Naik, Judith Rousseau, Trevor CampbellNeurIPS 2022 · 10 citations
- Dataset Distillation via the Wasserstein MetricHaoyang Liu, Yijiang Li, Tiancheng Xing, Peiran Wang et al.ICCV 2025 · 39 citations
- Large Scale Dataset Distillation with Domain ShiftNoel Loo, Alaa Maalouf, Ramin M. Hasani, Mathias Lechner et al.ICML 2024 · 9 citations
- General bounds on the quality of Bayesian coresetsTrevor CampbellNeurIPS 2024 · 3 citations
