Be Confident in What You Know: Bayesian Parameter Efficient Fine-Tuning of Vision Foundation Models
Deep Shankar Pandey, Spandan Pyakurel, Qi Yu
摘要
Large transformer-based foundation models have been commonly used as pre-trained models that can be adapted to different challenging datasets and settings with state-of-the-art generalization performance. Parameter efficient fine-tuning ( PEFT ) provides promising generalization performance in adaptation while incurring minimum computational overhead. However, adaptation of these foundation models through PEFT leads to accurate but severely underconfident models, especially in few-shot learning settings. Moreover, the adapted models lack accurate fine-grained uncertainty quantification capabilities limiting their broader applicability in critical domains. To fill out this critical gap, we develop a novel lightweight Bayesian Parameter Efficient Fine-Tuning (referred to as Bayesian-PEFT ) framework for large transformer-based foundation models. The framework integrates state-of-the-art PEFT techniques with two Bayesian components to address the under-confidence issue while ensuring reliable prediction under challenging few-shot settings. The first component performs base rate adjustment to strengthen the prior belief corresponding to the knowledge gained through pre-training, making the model more confident in its predictions; the second component builds an ev-idential ensemble that leverages belief regularization to ensure diversity among different ensemble components. Our thorough theoretical analysis justifies that the Bayesian components can ensure reliable and accurate few-shot adaptations with well-calibrated uncertainty quantification. Extensive experiments across diverse datasets, few-shot learning scenarios, and multiple PEFT techniques demonstrate the outstanding prediction and calibration performance by Bayesian-PEFT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Deep Evidential RegressionAlexander Amini, Wilko Schwarting, Ava Soleimany, Daniela RusNeurIPS 2020 · 被引用 777 次
相关 Paper
- Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual RecognitionZheda Mai, Ping Zhang, Cheng-Hao Tu, Hong-You Chen 等CVPR 2025
- Strong Baselines for Parameter-Efficient Few-Shot Fine-TuningSamyadeep Basu, Shell Xu Hu, Daniela Massiceti, Soheil FeiziAAAI 2024 · 被引用 54 次
- Weight-Space Learning for Certifiable Few-shot Transfer LearningFady Rezk, Royson Lee, Henry Gouk, Timothy Hospedales 等ICML 2026 · 被引用 1 次
- Multi-Head Adapter Routing for Cross-Task GeneralizationLucas Page-Caccia, Edoardo Maria Ponti, Zhan Su, Matheus Pereira 等NeurIPS 2023 · 被引用 38 次
- EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning In Vision TransformersWenwen Liao, Hang Ruan, Jianbo Yu, Bing Song 等AAAI 2026
