Variational Information Bottleneck for Effective Low-Resource Fine-Tuning
Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson
Abstract
While large-scale pretrained language models have obtained impressive results when fine-tuned on a wide variety of tasks, they still often suffer from overfitting in low-resource scenarios. Since such models are general-purpose feature extractors, many of these features are inevitably irrelevant for a given target task. We propose to use Variational Information Bottleneck (VIB) to suppress irrelevant features when fine-tuning on low-resource target tasks, and show that our method successfully reduces overfitting. Moreover, we show that our VIB model finds sentence representations that are more robust to biases in natural language inference datasets, and thereby obtains better generalization to out-of-domain datasets. Evaluation on seven low-resource datasets in different tasks shows that our method significantly improves transfer learning in low-resource scenarios, surpassing prior work. Moreover, it improves generalization on 13 out of 15 out-of-domain natural language inference benchmarks. Our code is publicly available in https://github.com/rabeehk/vibert .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers28
- Graph Structure Learning with Variational Information BottleneckQingyun Sun, Jianxin Li, Hao Peng, Jia Wu et al.AAAI 2022 · 224 citations
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningRunxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan et al.EMNLP 2021 · 129 citations
- Cross-Domain Recommendation to Cold-Start Users via Variational Information BottleneckJiangxia Cao, Jiawei Sheng, Xin Cong, Tingwen Liu et al.ICDE 2022 · 112 citations
- Supervising Model Attention with Human Explanations for Robust Natural Language InferenceJoe Stacey, Yonatan Belinkov, Marek ReiAAAI 2022 · 52 citations
- InfoPrompt: Information-Theoretic Soft Prompt Tuning for Natural Language UnderstandingJunda Wu, Tong Yu, Rui Wang, Zhao Song et al.NeurIPS 2023 · 48 citations
Builds on4
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 233 citations
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger et al.ICLR 2021 · 172 citations
- End-to-End Bias Mitigation by Modelling Biases in CorporaRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonACL 2020 · 136 citations
Related papers
- InfoBERT: Improving Robustness of Language Models from An Information Theoretic PerspectiveBoxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan et al.ICLR 2021 · 132 citations
- Information Retention via Learning Supplemental FeaturesZhipeng Xie, Yahe LiICLR 2024 · 1 citation
- Towards Robust Low-Resource Fine-Tuning with Multi-View Compressed RepresentationsLinlin Liu, Xingxuan Li, Megh Thakkar, Xin Li et al.ACL 2023 · 3 citations
- Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information FlowJiaqi Bai, Hongcheng Guo, Zhongyuan Peng, Jian Yang et al.AAAI 2025 · 7 citations
- Visual Instruction Bottleneck TuningChangdae Oh, Jiatong Li, Shawn Im, Sharon LiNeurIPS 2025 · 7 citations
