GIO: Gradient Information Optimization for Training Dataset Selection
Dante Everaert, Christopher Potts
Abstract
It is often advantageous to train models on a subset of the available train examples, because the examples are of variable quality or because one would like to train with fewer examples, without sacrificing performance. We present Gradient Information Optimization (GIO), a scalable, task-agnostic approach to this data selection problem that requires only a small set of (unlabeled) examples representing a target distribution. GIO begins from a natural, information-theoretic objective that is intractable in practice. Our contribution is in showing that it can be made highly scalable through a simple relaxation of the objective and a highly efficient implementation. In experiments with machine translation, spelling correction, and image recognition, we show that GIO delivers outstanding results with very small train sets. These findings are robust to different representation models and hyperparameters for GIO itself. GIO is task-and domain-agnostic and can be applied out-of-the-box to new datasets and domains. We open source a pip-installable implementation of the algorithm as "pip install grad-info-opt". 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2f5dd61-6974-4963-8f2c-011f45ccd156Cited by top-tier papers10
- Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMsFeiyang Kang, Hoang Anh Just, Yifan Sun, Himanshu Jahagirdar et al.ICLR 2024 · 39 citations
- Angles Don't Lie: Unlocking Training‑Efficient RL Through the Model's Own SignalsQinsi Wang, Jinghan Ke, Hancheng Ye, Yueqian Lin et al.NeurIPS 2025 · 16 citations
- How Reliable is Language Model Micro-Benchmarking?Gregory Yauney, Shahzaib Saqib Warraich, Swabha SwayamdiptaICLR 2026 · 7 citations
- Mapping Overlaps in Benchmarks through Perplexity in the WildSiyang Wu, Honglin Bao, Sida Li, Ari Holtzman et al.ICLR 2026 · 5 citations
- LAMDAS: LLM as an Implicit Classifier for Domain-specific Data SelectionJian Wu, Hang Yu, Bingchang Liu, Wenjie Yang et al.AAAI 2026 · 1 citation
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
Related papers
- FastClass: A Time-Efficient Approach to Weakly-Supervised Text ClassificationTingyu Xia, Yue Wang, Yuan Tian, Yi ChangEMNLP 2022 · 1 citation
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 74 citations
- Train on Validation (ToV): Fast data selection with applications to fine-tuningAyush Jain, Andrea Montanari, Eren SasogluICLR 2026 · 4 citations
- Disentangling the Roles of Representation and Selection in Data PruningYupei Du, Yingjin Song, Hugh Mee Wong, Daniil Ignatev et al.ACL 2025
- A Model-Agnostic Approach for Learning with Noisy Labels of Arbitrary DistributionsShuang Hao, Peng Li, Renzhi Wu, Xu ChuICDE 2022 · 2 citations
