Unifying and Optimizing Data Values for Selection via Sequential Decision-Making
Frank Hongliang Chi, Qiong Wu, Zhengyi Zhou, Jonathan Li, Emily Dodwell, Yao Ma
Abstract
Data selection has emerged as a crucial downstream application of data valuation, yet the theoretical foundations for using data values in selection remain underexplored. We reformulate data selection as a sequential decision-making problem where the optimal selection sequence arises from dynamic programming, and data values can be understood as encodings of this optimal sequence. This framework unifies and reinterprets existing methods like Data Shapley through the lens of approximate dynamic programming, revealing them as myopic linear approximations to the sequential problem. We further analyze how selection optimality degrades with utility curvature under submodularity, explaining when and why these approximations fail. To bridge theory and practice, we propose an efficient bipartite graph-based surrogate that preserves submodular structure while enabling scalable greedy selection with provable guarantees. Experiments on classical ML benchmarks and large-scale LLM fine-tuning data selection demonstrate substantial improvements over existing methods. Code is publicly available at https://github.com/ frankhlchi/SeqDataVal .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6880e07-bf1c-42b8-a66b-39dfb56c0218Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 236 citations
- DsDm: Model-Aware Dataset Selection with DatamodelsLogan Engstrom, Axel Feldmann, Aleksander MadryICML 2024 · 105 citations
- If You Like Shapley Then You'll Love the CoreTom Yan, Ariel D. ProcacciaAAAI 2021 · 85 citations
Related papers
- Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon et al.ICML 2024 · 24 citations
- Efficient Top-m Data Values Identification for Data SelectionXiaoqiang Lin, Xinyi Xu, See-Kiong Ng, Bryan Kian Hsiang LowICLR 2025
- EcoVal: An Efficient Data Valuation Framework for Machine LearningAyush K. Tarun, Vikram S. Chundawat, Murari Mandal, Hong Ming Tan et al.KDD 2024 · 3 citations
- Learning to Make Decisions via Submodular RegularizationAyya Alieva, Aiden Aceves, Jialin Song, Stephen Mayo et al.ICLR 2021 · 5 citations
- A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn’t)Nihal Nayak, Paula Rodriguez-Diaz, Neha Hulkund, Sara Beery et al.ICML 2026 · 2 citations
