Deep Reference Priors: What is the best way to pretrain a model?
Yansong Gao, Rahul Ramesh, Pratik Chaudhari
摘要
What is the best way to exploit extra data-be it unlabeled data from the same task, or labeled data from a related task-to learn a given task? This paper formalizes the question using the theory of reference priors. Reference priors are objective, uninformative Bayesian priors that maximize the mutual information between the task and the weights of the model. Such priors enable the task to maximally affect the Bayesian posterior, e.g., reference priors depend upon the number of samples available for learning the task and for very small sample sizes, the prior puts more probability mass on low-complexity models in the hypothesis space. This paper presents the first demonstration of reference priors for medium-scale deep networks and image-based data. We develop generalizations of reference priors and demonstrate applications to two problems. First, by using unlabeled data to compute the reference prior, we develop new Bayesian semi-supervised learning methods that remain effective even with very few samples per class. Second, by using labeled data from the source task to compute the reference prior, we develop a new pretraining method for transfer learning that allows data from the target task to maximally affect the Bayesian posterior. Empirical validation of these methods is conducted on image classification datasets. Code is available at https://github.com/grasp-lyrl/deep_reference_priors .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- A Baseline for Few-Shot Image ClassificationGuneet Singh Dhillon, Pratik Chaudhari, Avinash Ravichandran, Stefano SoattoICLR 2020 · 被引用 640 次
- Dash: Semi-Supervised Learning with Dynamic ThresholdingYi Xu, Lei Shang, Jinxing Ye, Qi Qian 等ICML 2021 · 被引用 287 次
- Model Zoo: A Growing Brain That Learns ContinuallyRahul Ramesh, Pratik ChaudhariICLR 2022 · 被引用 79 次
相关 Paper
- Pre-Train Your Loss: Easy Bayesian Transfer Learning with Informative PriorsRavid Shwartz-Ziv, Micah Goldblum, Hossein Souri, Sanyam Kapoor 等NeurIPS 2022 · 被引用 52 次
- Adversarial Training Helps Transfer Learning via Better RepresentationsZhun Deng, Linjun Zhang, Kailas Vodrahalli, Kenji Kawaguchi 等NeurIPS 2021 · 被引用 60 次
- TransMatch: A Transfer-Learning Scheme for Semi-Supervised Few-Shot LearningZhongjie Yu, Lin Chen, Zhongwei Cheng, Jiebo LuoCVPR 2020
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- How Does Semi-supervised Learning with Pseudo-labelers Work? A Case StudyYiwen Kou, Zixiang Chen, Yuan Cao, Quanquan GuICLR 2023
