LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection
Jian Wu, Hang Yu, Bingchang Liu, Wenjie Yang, Peng Di, Jianguo Li, Yue Zhang
摘要
Adapting large language models (LLMs) to specific domains often faces a critical bottleneck: the scarcity of high-quality, human-curated data. While large volumes of unchecked data are readily available, indiscriminately using them for fine-tuning risks introducing noise and degrading performance. Strategic data selection is thus crucial, requiring a method that is both accurate and efficient. Existing approaches, categorized as similarity-based and direct optimization methods, struggle to simultaneously achieve these goals. In this paper, we introduce LAMDAS (LLM As an iMplicit classifier for domain-specific DAta Selection), a novel approach that leverages the pre-trained LLM itself as an implicit classifier, thereby bypassing explicit feature engineering and computationally intensive optimization process. LAMDAS reframes data selection as a one-class classification problem, identifying candidate data that "belongs" to the target domain defined by a small reference dataset. Extensive experimental results demonstrate that LAMDAS not only exceeds the performance of full-data training using a fraction of the data but also outperforms nine state-of-the-art (SOTA) baselines under various scenarios. Furthermore, LAMDAS achieves the most compelling balance between performance gains and computational efficiency compared to all evaluated baselines. a a Code will be available upon publication.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
相关 Paper
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMsFeiyang Kang, Hoang Anh Just, Yifan Sun, Himanshu Jahagirdar 等ICLR 2024 · 被引用 39 次
- Unsupervised Domain Clusters in Pretrained Language ModelsRoee Aharoni, Yoav GoldbergACL 2020 · 被引用 13 次
- Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context LearningShaobo Wang, Xiangqi Jin, Ziming Wang, Jize Wang 等ACL 2025
- From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language ModelsLing Shi, Xinwei Wu, Xiaohu Zhao, Hao Wang 等ACL 2026
