TRACE: Trajectory-based Activation Change Estimation for Task-specific Data Selection
Ye He, Shangzhan Li, Yuxin Zhou, Qi Shi
Abstract
Task-specific data selection, which aims to identify the most relevant training instances from a large corpus to optimize performance on a target task, is a critical challenge in modern AI. Prevailing methods typically rely on either representation clustering or gradient-based influence estimation. However, these approaches have notable limitations. Representationbased methods rely on static features; they measure semantic proximity but are agnostic to the process of learning. Conversely, influence-based methods, while capturing optimization directions, often focus narrowly on aligning with the validation loss, which may not fully correlate with the desired capabilities. To address these issues, we propose TRACE, a novel algorithm that simultaneously considers data consistency in the optimization direction and representation space, and performs TRajectory-based Activation Change Estimation to select instruction data. Specifically, TRACE first performs a targeted weight update using the validation set. It then captures the optimization trajectory by calculating the change in neuron activations for each before and after this update. By selecting data whose activation change are most similar to those of the validation set, TRACE ensures alignment in both the representational and optimization domains. Our experiments demonstrate that TRACE outperforms baseline methods across various tasks, particularly in complex, data-scarce scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e9d0e99-d023-4d52-a71f-acfd19606c54Builds on30
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
Related papers
- Capturing the Temporal Dependence of Training Data InfluenceJiachen T. Wang, Dawn Song, James Zou, Prateek Mittal et al.ICLR 2025
- Tracing Training Progress: Dynamic Influence Based Selection for Active LearningTianjiao Wan, Kele Xu, Long Lan, Zijian Gao et al.ACM MM 2024 · 3 citations
- NICE Data Selection for Instruction Tuning in LLMs with Non-differentiable Evaluation MetricJingtan Wang, Xiaoqiang Lin, Rui Qiao, Pang Wei Koh et al.ICML 2025
- Understanding Data Influence in Reinforcement FinetuningHaoru Tan, Xiuzhe Wu, Sitong Wu, Shaofeng Zhang et al.NeurIPS 2025 · 4 citations
- Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal ActivationsDa Ma, Gonghu Shang, Zhi Chen, Libo Qin et al.NeurIPS 2025 · 6 citations
