Recursive PAC-Bayes: A Frequentist Approach to Sequential Prior Updates with No Information Loss
Yi-Shan Wu, Yijie Zhang, Badr-Eddine Chérief-Abdellatif, Yevgeny Seldin
摘要
PAC-Bayesian analysis is a frequentist framework for incorporating prior knowledge into learning. It was inspired by Bayesian learning, which allows sequential data processing and naturally turns posteriors from one processing step into priors for the next. However, despite two and a half decades of research, the ability to update priors sequentially without losing confidence information along the way remained elusive for PAC-Bayes. While PAC-Bayes allows construction of data-informed priors, the final confidence intervals depend only on the number of points that were not used for the construction of the prior, whereas confidence information in the prior, which is related to the number of points used to construct the prior, is lost. This limits the possibility and benefit of sequential prior updates, because the final bounds depend only on the size of the final batch. We present a novel and, in retrospect, surprisingly simple and powerful PAC-Bayesian procedure that allows sequential prior updates with no information loss. The procedure is based on a novel decomposition of the expected loss of randomized classifiers. The decomposition rewrites the loss of the posterior as an excess loss relative to a downscaled loss of the prior plus the downscaled loss of the prior, which is bounded recursively. As a side result, we also present a generalization of the split-kl and PAC-Bayes-split-kl inequalities to discrete random variables, which we use for bounding the excess losses, and which can be of independent interest. In empirical evaluation the new procedure significantly outperforms state-of-the-art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- How Tight Can PAC-Bayes be in the Small Data Regime?Andrew Y. K. Foong, Wessel P. Bruinsma, David R. Burt, Richard E. TurnerNeurIPS 2021 · 被引用 28 次
- Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority VoteYi-Shan Wu, Andrés R. Masegosa, Stephan Sloth Lorenzen, Christian Igel 等NeurIPS 2021 · 被引用 14 次
- Split-kl and PAC-Bayes-split-kl Inequalities for Ternary Random VariablesYi-Shan Wu, Yevgeny SeldinNeurIPS 2022
相关 Paper
- An Exact Characterization of the Generalization Error for the Gibbs AlgorithmGholamali Aminian, Yuheng Bu, Laura Toni, Miguel R. D. Rodrigues 等NeurIPS 2021 · 被引用 75 次
- PAC-Bayes Analysis Beyond the Usual BoundsOmar Rivasplata, Ilja Kuzborskij, Csaba Szepesvári, John Shawe-TaylorNeurIPS 2020 · 被引用 101 次
- PAC-Bayes Learning Bounds for Sample-Dependent PriorsPranjal Awasthi, Satyen Kale, Stefani Karp, Mehryar MohriNeurIPS 2020 · 被引用 6 次
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 被引用 10 次
- Online PAC-Bayes LearningMaxime Haddouche, Benjamin GuedjNeurIPS 2022 · 被引用 33 次
