Selecting Informative Contexts Improves Language Model Fine-tuning
Richard J. Antonello, Nicole Beckage, Javier Turek, Alexander Huth
摘要
Language model fine-tuning is essential for modern natural language processing, but is computationally expensive and timeconsuming. Further, the effectiveness of finetuning is limited by the inclusion of training examples that negatively affect performance. Here we present a general fine-tuning method that we call information gain filtration for improving the overall training efficiency and final performance of language model fine-tuning. We define the information gain of an example as the improvement on a validation metric after training on that example. A secondary learner is then trained to approximate this quantity. During fine-tuning, this learner selects informative examples and skips uninformative ones. We show that our method has consistent improvement across datasets, finetuning tasks, and language model architectures. For example, we achieve a median perplexity of 54.0 on a books dataset compared to 57.3 for standard fine-tuning. We present statistical evidence that offers insight into the improvements of our method over standard finetuning. The generality of our method leads us to propose a new paradigm for language model fine-tuning -we encourage researchers to release pretrained secondary learners on common corpora to promote efficient and effective fine-tuning, thereby improving the performance and reducing the overall energy footprint of language model fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Leveraging Similar Users for Personalized Language Modeling with Limited DataCharles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas 等ACL 2022 · 被引用 37 次
- Efficient Data Selection at Scale via Influence DistillationMahdi Nikdan, Vincent Cohen-Addad, Dan Alistarh, Vahab MirrokniNeurIPS 2025 · 被引用 15 次
- Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline MethodsWanru Zhao, Yihong Chen, Yuzhi Tang, Wentao Ma 等ICLR 2026 · 被引用 4 次
- GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization GeometryGuanghui Min, Tianhao Huang, Ke Wan, Chen ChenICML 2026 · 被引用 3 次
- Compute-Constrained Data SelectionJunjie Oscar Yin, Alexander M. RushICLR 2025
它引用的顶会 Paper3
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 被引用 448 次
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 被引用 233 次
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger 等ICLR 2021 · 被引用 172 次
相关 Paper
- FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information GainRohan Deb, Kiran Koshy Thekumparampil, Kousha Kalantari, Gaurush Hiranandani 等ICML 2025
- Double-Filter: Efficient Fine-tuning of Pre-trained Vision-Language Models via Patch&Layer FilteringYaoqin He, Junchen Fu, Kaiwen Zheng, Songpei Xu 等ICML 2025
- Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-TuningMing Li, Yong Zhang, Shwai He, Zhitao Li 等ACL 2024 · 被引用 16 次
- DELIFT: Data Efficient Language model Instruction Fine-TuningIshika Agarwal, Krishnateja Killamsetty, Lucian Popa, Marina DanilevskyICLR 2025
- Towards Green AI in Fine-tuning Large Language Models via Adaptive BackpropagationKai Huang, Hanyun Yin, Heng Huang, Wei GaoICLR 2024 · 被引用 21 次
