Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
Jinlong Pang, Na Di, Zhaowei Zhu, Jiaheng Wei, Hao Cheng, Chen Qian, Yang Liu
Abstract
Recent studies show that in supervised fine-tuning (SFT) of large language models (LLMs), data quality matters more than quantity. While most data cleaning methods concentrate on filtering entire samples, the quality of individual tokens within a sample can vary significantly. After pretraining, even in high-quality samples, patterns or phrases that are not task-related can be redundant, uninformative, or even harmful. Continuing to fine-tune on these patterns may offer limited benefit and even degrade downstream task performance. In this paper, we investigate token quality from a noisy-label perspective and propose a generic token cleaning pipeline for SFT tasks. Our method filters out uninformative tokens while preserving those carrying key task-specific information. Specifically, we first evaluate token quality by examining the influence of model updates on each token, then apply a thresholdbased separation. The token influence can be measured in a single pass with a fixed reference model or iteratively with self-evolving reference models. The benefits and limitations of both methods are analyzed theoretically by error upper bounds. Extensive experiments show that our framework consistently improves downstream performance. Code is available at https://github.com/ UCSC-REAL/TokenCleaning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenCVPR 2026 · 18 citations
- T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction TuningYanjun Fu, Faisal Hamman, Sanghamitra DuttaNeurIPS 2025 · 15 citations
- IF-Guide: Influence Function-Guided Detoxification of LLMsZachary Coalson, Juhan Bae, Nicholas Carlini, Sanghyun HongNeurIPS 2025 · 9 citations
- Towards Understanding Valuable Preference Data for Large Language Model AlignmentZizhuo Zhang, Qizhou Wang, Shanshan Ye, Jianing Zhu et al.ICLR 2026 · 6 citations
- Token-level Data Selection for Safe LLM Fine-tuningYanping Li, Zhening Liu, Zijian Li, Zehong Lin et al.ICLR 2026 · 4 citations
Builds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction TuningWei Liu, Weihao Zeng, Keqing He, Yong Jiang et al.ICLR 2024 · 369 citations
Related papers
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuningXiaohan Qin, Victor Wang, Ning Liao, Cancheng Zhang et al.ICLR 2026 · 3 citations
- Explainable Token-level Noise Filtering for LLM Fine-tuning DatasetsYuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu et al.ICLR 2026 · 1 citation
- SELECting over Tokens: Curating Pre-training Data at Scale via Token ClassificationXin Tong, Weidong Zhang, Jiaang Li, Haibin Chen et al.ACL 2026
- Make Every Example Count: On the Stability and Utility of Self-Influence for Learning from Noisy NLP DatasetsIrina Bejan, Artem Sokolov, Katja FilippovaEMNLP 2023 · 7 citations
- Clipping Low-Probability Tokens in SFT Yields a Generalizable Initialization for RLTian-Shuo Liu, Chengxing Jia, Haoyu Liu, Pengyuan Wang et al.ICML 2026
