Selecting Informative Contexts Improves Language Model Fine-tuning
Richard J. Antonello, Nicole Beckage, Javier Turek, Alexander Huth
Abstract
Language model fine-tuning is essential for modern natural language processing, but is computationally expensive and timeconsuming. Further, the effectiveness of finetuning is limited by the inclusion of training examples that negatively affect performance. Here we present a general fine-tuning method that we call information gain filtration for improving the overall training efficiency and final performance of language model fine-tuning. We define the information gain of an example as the improvement on a validation metric after training on that example. A secondary learner is then trained to approximate this quantity. During fine-tuning, this learner selects informative examples and skips uninformative ones. We show that our method has consistent improvement across datasets, finetuning tasks, and language model architectures. For example, we achieve a median perplexity of 54.0 on a books dataset compared to 57.3 for standard fine-tuning. We present statistical evidence that offers insight into the improvements of our method over standard finetuning. The generality of our method leads us to propose a new paradigm for language model fine-tuning -we encourage researchers to release pretrained secondary learners on common corpora to promote efficient and effective fine-tuning, thereby improving the performance and reducing the overall energy footprint of language model fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9efeda2a-5728-4bfe-990e-1716967d9ae5Cited by top-tier papers6
- Leveraging Similar Users for Personalized Language Modeling with Limited DataCharles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas et al.ACL 2022 · 37 citations
- Efficient Data Selection at Scale via Influence DistillationMahdi Nikdan, Vincent Cohen-Addad, Dan Alistarh, Vahab MirrokniNeurIPS 2025 · 15 citations
- Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline MethodsWanru Zhao, Yihong Chen, Yuzhi Tang, Wentao Ma et al.ICLR 2026 · 4 citations
- GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization GeometryGuanghui Min, Tianhao Huang, Ke Wan, Chen ChenICML 2026 · 3 citations
- Compute-Constrained Data SelectionJunjie Oscar Yin, Alexander M. RushICLR 2025
Builds on3
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 233 citations
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger et al.ICLR 2021 · 172 citations
Related papers
- FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information GainRohan Deb, Kiran Koshy Thekumparampil, Kousha Kalantari, Gaurush Hiranandani et al.ICML 2025
- Double-Filter: Efficient Fine-tuning of Pre-trained Vision-Language Models via Patch&Layer FilteringYaoqin He, Junchen Fu, Kaiwen Zheng, Songpei Xu et al.ICML 2025
- Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-TuningMing Li, Yong Zhang, Shwai He, Zhitao Li et al.ACL 2024 · 16 citations
- DELIFT: Data Efficient Language model Instruction Fine-TuningIshika Agarwal, Krishnateja Killamsetty, Lucian Popa, Marina DanilevskyICLR 2025
- Towards Green AI in Fine-tuning Large Language Models via Adaptive BackpropagationKai Huang, Hanyun Yin, Heng Huang, Wei GaoICLR 2024 · 21 citations
