Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
Ming Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, Tianyi Zhou
Abstract
Instruction tuning is critical to improve LLMs but usually suffers from low-quality and redundant data. Data filtering for instruction tuning has proved important in improving both the efficiency and performance of the tuning process. But it also leads to extra cost and computation due to the involvement of LLMs in this process. To reduce the filtering cost, we study Superfiltering: Can we use a smaller and weaker model to select data for finetuning a larger and stronger model? Despite the performance gap between weak and strong language models, we find their highly consistent capability to perceive instruction difficulty and data selection results. This enables us to use a much smaller and more efficient model to filter the instruction data used to train a larger language model. Not only does it largely speed up the data filtering, but the filtered-data-finetuned LLM achieves even better performance on standard benchmarks. Extensive experiments validate the efficacy and efficiency of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69953f2e-3d2f-4b0c-8fe1-cd6de95ddbc1Cited by top-tier papers53
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-ImprovementXiyao Wang, Zhengyuan Yang, Chao Feng, Hongjin Lu et al.NeurIPS 2025 · 158 citations
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 59 citations
- Star-Agents: Automatic Data Optimization with LLM Agents for Instruction TuningHang Zhou, Yehui Tang, Haochen Qin, Yujie Yang et al.NeurIPS 2024 · 21 citations
- LEAD: Iterative Data Selection for Efficient LLM Instruction TuningXiaotian Lin, Yanlin Qi, Yizhang Zhu, Themis Palpanas et al.VLDB 2026 · 16 citations
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsPeng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen et al.NeurIPS 2025 · 16 citations
Builds on18
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
Related papers
- CPQS-Tuning: A Model Self-Perception-Based Data Filtering Algorithm for Efficient Instruction Fine-TuningYI Ren, Yanhui Li, Tianyi Zhang, Diandong LiuICLR 2026
- Importance-Aware Data Selection for Efficient LLM Instruction TuningTingyu Jiang, Shen Li, Yiyao Song, Lan Zhang et al.AAAI 2026 · 5 citations
- DELIFT: Data Efficient Language model Instruction Fine-TuningIshika Agarwal, Krishnateja Killamsetty, Lucian Popa, Marina DanilevskyICLR 2025
- BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMsChaoyuan Shen, Chi Zhang, Chengliang Chai, Jiacheng Wang et al.VLDB 2026 · 2 citations
- Compute-Constrained Data SelectionJunjie Oscar Yin, Alexander M. RushICLR 2025
