Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations
Da Ma, Gonghu Shang, Zhi Chen, Libo Qin, Yijie Luo, Hongshen Xu, Lei Pan, Shuai Fan, Kai Yu, Lu Chen
Abstract
Instruction tuning improves the ability of large language models (LLMs) to follow diverse human instructions, but achieving strong performance on specific target tasks remains challenging. A critical bottleneck is selecting the most relevant data to maximize task-specific performance. Existing data selection approaches include unstable influence-based methods and more stable distribution alignment methods, the latter of which critically rely on the underlying sample representation. In practice, most distribution alignment methods, from shallow features (e.g., BM25) to neural embeddings (e.g., BGE, LLM2Vec), may fail to capture how the model internally processes samples. To bridge this gap, we adopt a model-centric strategy in which each sample is represented by its neuronal activation pattern in the model, directly reflecting internal computation. However, directly using raw neuron activations leads to spurious similarity between unrelated samples due to neuron polysemanticity, where a single neuron may respond to multiple, unrelated concepts. To address this, we employ sparse autoencoders to disentangle polysemantic activations into sparse, monosemantic representations, and introduce a dedicated similarity metric for this space to better identify task-relevant data. Comprehensive experiments across multiple instruction datasets, models, tasks, and selection ratios show that our approach consistently outperforms existing data selection baselines in both stability and task-specific performance 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 726822c1-893b-428f-99cd-76503c0c4ff9Cited by top-tier papers2
- TRACE: Trajectory-based Activation Change Estimation for Task-specific Data SelectionYe He, Shangzhan Li, Yuxin Zhou, Qi ShiAAAI 2026
- Matched Data, Better Models: Target Aligned Data Filtering with Sparse AutoencodersArnav Das, Gantavya Bhatt, Sahil Verma, Yiping Wang et al.ICLR 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
Related papers
- Task-Aware Data Selection via Proxy-Label Enhanced Distribution Matching for LLM FinetuningHao Cheng, Rui Zhang, Ling Li, Na Di et al.ICLR 2026
- A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn’t)Nihal Nayak, Paula Rodriguez-Diaz, Neha Hulkund, Sara Beery et al.ICML 2026 · 2 citations
- Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMsXinwei Wu, Heng Liu, Xiaohu Zhao, Yuqi Ren et al.AAAI 2026 · 2 citations
- Neuron-Aware Data Selection in Instruction Tuning for Large Language ModelsXin Chen, Junchao Wu, Shu Yang, Runzhe Zhan et al.ICLR 2026 · 2 citations
- Wasserstein Distances, Neuronal Entanglement, and SparsityShashata Sawmya, Linghao Kong, Ilia Markov, Dan Alistarh et al.ICLR 2025
