Mining Useful General Data for Low-Resource Domain Adaptation
Pingjie Wang, Hongcheng Liu, Yusheng Liao, Ziqing Fan, Yaxin Du, shuo tang, Yanfeng Wang, Yu Wang
Abstract
Adapting large language models (LLMs) to low-resource domains remains challenging due to the scarcity of domain-specific data. While in-domain data is limited, there exists a vast amount of general-domain data that shares similar question–answer formats and reasoning patterns with domain tasks. This observation raises an important question: can useful general-domain data be mined to improve low-resource domain adaptation? Our initial findings show that general-domain chain-of-thought data contains useful auxiliary signals for domain adaptation, even without careful selection. This observation motivates a new paradigm for domain adaptation beyond exclusive reliance on domain-specific data. To systematically identify the most beneficial general-domain samples, we propose NTK-Selector, motivated by the Neural Tangent Kernel’s ability to capture alignment in training dynamics. Since directly applying NTK to pretrained LLMs is impractical, we introduce a Jacobian-free NTK approximation and empirically demonstrate stable NTK-like behavior during fine-tuning. Extensive experiments across medical, financial, legal, and psychological domains demonstrate that NTK-Selector consistently outperforms domain-only fine-tuning and existing data selection baselines. In particular, NTK-Selector achieves gains of +8.7 and +5.1 points on Llama3-8B-Instruct and Qwen3-8B, respectively, compared to only +0.8 and +0.9 points from domain-only fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efb9b421-973e-4351-8db0-2521382e1140Builds on17
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li et al.NeurIPS 2021 · 303 citations
- AlpaGasus: Training a Better Alpaca with Fewer DataLichang Chen, Shiyang Li, Jun Yan, Hai Wang et al.ICLR 2024 · 295 citations
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen et al.ICML 2023 · 111 citations
Related papers
- LensLLM: Unveiling Fine-Tuning Dynamics for LLM SelectionXinyue Zeng, Haohui Wang, Junhong Lin, Jun Wu et al.ICML 2025
- Understanding Linear Probing then Fine-tuning Language Models from NTK PerspectiveAkiyoshi Tomihari, Issei SatoNeurIPS 2024 · 30 citations
- Linearization Explains Fine-Tuning in Large Language ModelsZahra Rahimi Afzal, Tara Esmaeilbeig, Mojtaba Soltanalian, Mesrob I. OhannessianNeurIPS 2025 · 5 citations
- LoRA Training in the NTK Regime has No Spurious Local MinimaUijeong Jang, Jason D. Lee, Ernest K. RyuICML 2024 · 41 citations
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 4 citations
