Convex Dataset Valuation for Post-Training
Siqi Zeng, Christopher Jung, Rui Li, Zhe Kang, Ming Li, Nima Noorshams, Zhigang Wang, Fuchun Peng, Han Zhao, Xue Feng
摘要
Improving LLM performance on downstream tasks sometimes requires leveraging auxiliary datasets during post-training. In practice, however, developers face constraints on compute, labeling, and licensing costs that preclude using all available data, necessitating principled datasetlevel selection. These constraints are increasingly shaped by dataset marketplaces, where data acquisition is governed by budgets and negotiation. We study dataset valuation as a subset selection problem during LLM post-training. Our goal is to identify and weight auxiliary datasets so as to maximize target task performance given constrained budgets. We first show that commonly used gradient alignment scores provide a reasonable yet incomplete valuation signal, as they ignore redundancy among datasets. To address this, we propose a scalable convex dataset-level valuation method based on kernel mean matching (KMM) in gradient space, which jointly accounts for alignment with the target task and redundancy across auxiliary datasets. Through extensive experiments across diverse post-training settings and tasks, we show that our approach consistently outperforms existing valuation baselines, achieving stronger performance with low computational overhead. Our results position dataset valuation as a practical decision tool for posttraining data selection in market-constrained large language model settings. The code is available at https://github.com/uiuctml/conve x_data_valuation .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
相关 Paper
- Automatic Auxiliary Task Selection and Adaptive Weighting Boost Molecular Property PredictionZhiqiang Zhong, Davide MottinNeurIPS 2025 · 被引用 3 次
- Data Selection for LLM Alignment Using Fine-Grained PreferencesJia Zhang, Yao Liu, Chen-Xi Zhang, Yi Liu 等ICLR 2026 · 被引用 1 次
- CondenseLM: LLMs-driven Text Dataset Condensation via Reward MatchingCheng Shen, Yew-Soon Ong, Joey Tianyi ZhouEMNLP 2025
- Minibatch selection for Language Models via Partition Matroid Constrained Gradient MatchingPrayas Agrawal, Prateek Chanda, Ishita Khatri, Ganesh Ramakrishnan 等ICML 2026
- DsDm: Model-Aware Dataset Selection with DatamodelsLogan Engstrom, Axel Feldmann, Aleksander MadryICML 2024 · 被引用 105 次
