Mining Useful General Data for Low-Resource Domain Adaptation
Pingjie Wang, Hongcheng Liu, Yusheng Liao, Ziqing Fan, Yaxin Du, shuo tang, Yanfeng Wang, Yu Wang
摘要
Adapting large language models (LLMs) to low-resource domains remains challenging due to the scarcity of domain-specific data. While in-domain data is limited, there exists a vast amount of general-domain data that shares similar question–answer formats and reasoning patterns with domain tasks. This observation raises an important question: can useful general-domain data be mined to improve low-resource domain adaptation? Our initial findings show that general-domain chain-of-thought data contains useful auxiliary signals for domain adaptation, even without careful selection. This observation motivates a new paradigm for domain adaptation beyond exclusive reliance on domain-specific data. To systematically identify the most beneficial general-domain samples, we propose NTK-Selector, motivated by the Neural Tangent Kernel’s ability to capture alignment in training dynamics. Since directly applying NTK to pretrained LLMs is impractical, we introduce a Jacobian-free NTK approximation and empirically demonstrate stable NTK-like behavior during fine-tuning. Extensive experiments across medical, financial, legal, and psychological domains demonstrate that NTK-Selector consistently outperforms domain-only fine-tuning and existing data selection baselines. In particular, NTK-Selector achieves gains of +8.7 and +5.1 points on Llama3-8B-Instruct and Qwen3-8B, respectively, compared to only +0.8 and +0.9 points from domain-only fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li 等NeurIPS 2021 · 被引用 303 次
- AlpaGasus: Training a Better Alpaca with Fewer DataLichang Chen, Shiyang Li, Jun Yan, Hai Wang 等ICLR 2024 · 被引用 295 次
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen 等ICML 2023 · 被引用 111 次
相关 Paper
- LensLLM: Unveiling Fine-Tuning Dynamics for LLM SelectionXinyue Zeng, Haohui Wang, Junhong Lin, Jun Wu 等ICML 2025
- Understanding Linear Probing then Fine-tuning Language Models from NTK PerspectiveAkiyoshi Tomihari, Issei SatoNeurIPS 2024 · 被引用 30 次
- Linearization Explains Fine-Tuning in Large Language ModelsZahra Rahimi Afzal, Tara Esmaeilbeig, Mojtaba Soltanalian, Mesrob I. OhannessianNeurIPS 2025 · 被引用 5 次
- LoRA Training in the NTK Regime has No Spurious Local MinimaUijeong Jang, Jason D. Lee, Ernest K. RyuICML 2024 · 被引用 41 次
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 被引用 4 次
