On Divergence Measures for Bayesian Pseudocoresets
Balhae Kim, Jungwon Choi, Seanie Lee, Yoonho Lee, Jung-Woo Ha, Juho Lee
摘要
A Bayesian pseudocoreset is a small synthetic dataset for which the posterior over parameters approximates that of the original dataset. While promising, the scalability of Bayesian pseudocoresets is not yet validated in realistic problems such as image classification with deep neural networks. On the other hand, dataset distillation methods similarly construct a small dataset such that the optimization using the synthetic dataset converges to a solution with performance competitive with optimization using full data. Although dataset distillation has been empirically verified in large-scale settings, the framework is restricted to point estimates, and their adaptation to Bayesian inference has not been explored. This paper casts two representative dataset distillation algorithms as approximations to methods for constructing pseudocoresets by minimizing specific divergence measures: reverse KL divergence and Wasserstein distance. Furthermore, we provide a unifying view of such divergence measures in Bayesian pseudocoreset construction. Finally, we propose a novel Bayesian pseudocoreset algorithm based on minimizing forward KL divergence. Our empirical results demonstrate that the pseudocoresets constructed from these methods reflect the true posterior even in high-dimensional Bayesian inference problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Multisize Dataset CondensationYang He, Lingao Xiao, Joey Tianyi Zhou, Ivor W. TsangICLR 2024 · 被引用 22 次
- Low-Rank Similarity Mining for Multimodal Dataset DistillationYue Xu, Zhilin Lin, Yusong Qiu, Cewu Lu 等ICML 2024 · 被引用 14 次
- Function Space Bayesian Pseudocoreset for Bayesian Neural NetworksBalhae Kim, Hyungi Lee, Juho LeeNeurIPS 2023 · 被引用 3 次
- CHESS: Chebyshev Spectral Synthesis for Trajectory CondensationRuituo Wu, Hongyu Zhang, Qiang Wang, Jiawei Du 等ICML 2026
- Variational Bayesian Pseudo-CoresetHyungi Lee, Seungyoo Lee, Juho LeeICLR 2025
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski 等ICML 2020 · 被引用 409 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
相关 Paper
- Bayesian PseudocoresetsDionysis Manousakas, Zuheng Xu, Cecilia Mascolo, Trevor CampbellNeurIPS 2020 · 被引用 35 次
- Fast Bayesian Coresets via Subsampling and Quasi-Newton RefinementCian Naik, Judith Rousseau, Trevor CampbellNeurIPS 2022 · 被引用 10 次
- Dataset Distillation via the Wasserstein MetricHaoyang Liu, Yijiang Li, Tiancheng Xing, Peiran Wang 等ICCV 2025 · 被引用 39 次
- Large Scale Dataset Distillation with Domain ShiftNoel Loo, Alaa Maalouf, Ramin M. Hasani, Mathias Lechner 等ICML 2024 · 被引用 9 次
- General bounds on the quality of Bayesian coresetsTrevor CampbellNeurIPS 2024 · 被引用 3 次
