SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
Liangxin Liu, Xuebo Liu, Derek F. Wong, Dongfang Li, Ziyi Wang, Baotian Hu, Min Zhang
摘要
Instruction tuning (IT) is crucial to tailoring large language models (LLMs) towards human-centric interactions. Recent advancements have shown that the careful selection of a small, high-quality subset of IT data can significantly enhance the performance of LLMs. Despite this, common approaches often rely on additional models or data, which increases costs and limits widespread adoption. In this work, we propose a novel approach, termed SelectIT, that capitalizes on the foundational capabilities of the LLM itself. Specifically, we exploit the intrinsic uncertainty present in LLMs to more effectively select high-quality IT data, without the need for extra resources. Furthermore, we introduce a curated IT dataset, the Selective Alpaca, created by applying SelectIT to the Alpaca-GPT4 dataset. Empirical results demonstrate that IT using Selective Alpaca leads to substantial model ability enhancement. The robustness of SelectIT has also been corroborated in various foundation models and domain-specific tasks. Our findings suggest that longer and more computationally intensive IT data may serve as superior sources of IT, offering valuable insights for future research in this area. Data, code, and scripts are freely available at https://github.com/Blue-Raincoat/SelectIT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget ReallocationFeng Xiong, Hongling Xu, Yifei Wang, Runxi Cheng 等EMNLP 2025 · 被引用 18 次
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise GradientsMing Li, Yanhong Li, Ziyue Li, Tianyi ZhouACL 2026 · 被引用 9 次
- IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model AlignmentChenlin Ming, Chendi Qu, Qizhi Pei, Zhuoshi Pan 等ICLR 2026 · 被引用 8 次
- SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model TrainingPowei Chang, Jinpeng Zhang, Bowen Chen, Chenyu Wang 等ICLR 2026 · 被引用 5 次
- REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient ReasoningHexuan Deng, Wenxiang Jiao, Xuebo Liu, Jun Rao 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
相关 Paper
- Neuron-Aware Data Selection in Instruction Tuning for Large Language ModelsXin Chen, Junchao Wu, Shu Yang, Runzhe Zhan 等ICLR 2026 · 被引用 2 次
- Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality EstimationYuan Ge, Yilun Liu, Chi Hu, Weibin Meng 等EMNLP 2024 · 被引用 10 次
- From Selection to Refinement: Iterative Optimization for Instruction DataHang Hu, Ziyan Liu, Rujie Wen, Ruihui Hou 等ACL 2026
- JI2S: Joint Influence-Aware Instruction Data Selection for Efficient Fine-TuningJingyu Wei, Bo Liu, Tianjiao Wan, Baoyun Peng 等EMNLP 2025
- Priority on High-Quality: Selecting Instruction Data via Consistency Verification of Noise InjectionHong Zhang, Feng Zhao, Ruilin Zhao, Cheng Yan 等EMNLP 2025
