Optimizing Data Usage via Differentiable Rewards
Xinyi Wang, Hieu Pham, Paul Michel, Antonios Anastasopoulos, Jaime G. Carbonell, Graham Neubig
Abstract
To acquire a new skill, humans learn better and faster if a tutor informs them of how much attention they should pay to particular content or practice problems based on their current knowledge level. Similarly, a machine learning model could potentially be trained better if data is presented in a way that adapts to its current learning state. In this paper, we examine the problem of training an adaptive scorer that weights data instances to maximally benefit learning. Training such as scorer efficiently is a challenging problem; in order to precisely quantify the effect of a data instance on the final model, a naive approach would require completing the entire training process and observing final performance. We propose an efficient alternative -Differentiable Data Selection (DDS) -that formulates a scorer as a learnable function of the training data that can be efficiently updated along with the main model being trained. Specifically, DDS updates the scorer with an intuitive reward signal: it should up-weigh the data that has a similar gradient with a development set upon which we would finally like to perform well. Without significant computing overhead, DDS delivers consistent improvements over several strong baselines on two very different tasks of machine translation and image classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
- DsDm: Model-Aware Dataset Selection with DatamodelsLogan Engstrom, Axel Feldmann, Aleksander MadryICML 2024 · 105 citations
- Curriculum Disentangled Recommendation with Noisy Multi-feedbackHong Chen, Yudong Chen, Xin Wang, Ruobing Xie et al.NeurIPS 2021 · 88 citations
Related papers
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 74 citations
- LLM Data Selection and Utilization via Dynamic Bi-level OptimizationYang Yu, Kai Han, Hang Zhou, Yehui Tang et al.ICML 2025
- Learning to Reweight with Deep InteractionsYang Fan, Yingce Xia, Lijun Wu, Shufang Xie et al.AAAI 2021 · 6 citations
- DataRater: Meta-Learned Dataset CurationDan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf et al.NeurIPS 2025 · 17 citations
- Tracing Training Progress: Dynamic Influence Based Selection for Active LearningTianjiao Wan, Kele Xu, Long Lan, Zijian Gao et al.ACM MM 2024 · 3 citations
