Selection of LLM Fine-Tuning Data Based on Orthogonal Rules
Xiaomin Li, Mingye Gao, Zhiwei Zhang, Chang Yue, Hong Hu
Abstract
High-quality training data is critical to the performance of large language models (LLMs). Recent work has explored using LLMs to rate and select data based on a small set of human-designed criteria (rules), but these approaches often rely heavily on heuristics, lack principled metrics for rule evaluation, and generalize poorly to new tasks. We propose a novel rule-based data selection framework that introduces a metric based on the orthogonality of rule score vectors to evaluate and select complementary rules. Our automated pipeline first uses LLMs to generate diverse rules covering multiple aspects of data quality, then rates samples according to these rules and applies the determinantal point process (DPP) to select the most independent rules. These rules are then used to score the full dataset, and high-scoring samples are selected for downstream tasks such as LLM fine-tuning. We evaluate our framework in two experiment setups: (1) alignment with ground-truth ratings and (2) performance of LLMs fine-tuned on the selected data. Experiments across IMDB, Medical, Math, and Code domains demonstrate that our DPP-based rule selection consistently improves both rating accuracy and downstream model performance over strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8b36201-72d0-4a46-b9e4-7c22c2340c1fBuilds on10
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang et al.NeurIPS 2023 · 463 citations
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
Related papers
- Post-training Large Language Models for Diverse High-Quality ResponsesYilei Chen, Souradip Chakraborty, Lorenz Wolf, Ioannis Paschalidis et al.ICLR 2026 · 20 citations
- CritiQ: Mining Data Quality Criteria from Human PreferencesHonglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang et al.ACL 2025
- VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMsAvinash Amballa, Yashas Malur Saidutta, Chi-Heng Lin, Vivek Kulkarni et al.ACL 2026
- R-Select: A Robust Multi-Metric Data Selection Approach for Fine-Tuning Large Language ModelsXin Gao, Xiaoyang Wang, Yun Zhu, Zheng Liu et al.KDD 2026
- ActiveDPO: Active Direct Preference Optimization for Sample-Efficient AlignmentXiaoqiang Lin, Arun Verma, Zhongxiang Dai, Daniela Rus et al.ICLR 2026 · 12 citations
