GIO: Gradient Information Optimization for Training Dataset Selection
Dante Everaert, Christopher Potts
摘要
It is often advantageous to train models on a subset of the available train examples, because the examples are of variable quality or because one would like to train with fewer examples, without sacrificing performance. We present Gradient Information Optimization (GIO), a scalable, task-agnostic approach to this data selection problem that requires only a small set of (unlabeled) examples representing a target distribution. GIO begins from a natural, information-theoretic objective that is intractable in practice. Our contribution is in showing that it can be made highly scalable through a simple relaxation of the objective and a highly efficient implementation. In experiments with machine translation, spelling correction, and image recognition, we show that GIO delivers outstanding results with very small train sets. These findings are robust to different representation models and hyperparameters for GIO itself. GIO is task-and domain-agnostic and can be applied out-of-the-box to new datasets and domains. We open source a pip-installable implementation of the algorithm as "pip install grad-info-opt". 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMsFeiyang Kang, Hoang Anh Just, Yifan Sun, Himanshu Jahagirdar 等ICLR 2024 · 被引用 39 次
- Angles Don't Lie: Unlocking Training‑Efficient RL Through the Model's Own SignalsQinsi Wang, Jinghan Ke, Hancheng Ye, Yueqian Lin 等NeurIPS 2025 · 被引用 16 次
- How Reliable is Language Model Micro-Benchmarking?Gregory Yauney, Shahzaib Saqib Warraich, Swabha SwayamdiptaICLR 2026 · 被引用 7 次
- Mapping Overlaps in Benchmarks through Perplexity in the WildSiyang Wu, Honglin Bao, Sida Li, Ari Holtzman 等ICLR 2026 · 被引用 5 次
- LAMDAS: LLM as an Implicit Classifier for Domain-specific Data SelectionJian Wu, Hang Yu, Bingchang Liu, Wenjie Yang 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
相关 Paper
- FastClass: A Time-Efficient Approach to Weakly-Supervised Text ClassificationTingyu Xia, Yue Wang, Yuan Tian, Yi ChangEMNLP 2022 · 被引用 1 次
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 被引用 74 次
- Train on Validation (ToV): Fast data selection with applications to fine-tuningAyush Jain, Andrea Montanari, Eren SasogluICLR 2026 · 被引用 4 次
- Disentangling the Roles of Representation and Selection in Data PruningYupei Du, Yingjin Song, Hugh Mee Wong, Daniil Ignatev 等ACL 2025
- A Model-Agnostic Approach for Learning with Noisy Labels of Arbitrary DistributionsShuang Hao, Peng Li, Renzhi Wu, Xu ChuICDE 2022 · 被引用 2 次
