TRACE: Trajectory-based Activation Change Estimation for Task-specific Data Selection
Ye He, Shangzhan Li, Yuxin Zhou, Qi Shi
摘要
Task-specific data selection, which aims to identify the most relevant training instances from a large corpus to optimize performance on a target task, is a critical challenge in modern AI. Prevailing methods typically rely on either representation clustering or gradient-based influence estimation. However, these approaches have notable limitations. Representationbased methods rely on static features; they measure semantic proximity but are agnostic to the process of learning. Conversely, influence-based methods, while capturing optimization directions, often focus narrowly on aligning with the validation loss, which may not fully correlate with the desired capabilities. To address these issues, we propose TRACE, a novel algorithm that simultaneously considers data consistency in the optimization direction and representation space, and performs TRajectory-based Activation Change Estimation to select instruction data. Specifically, TRACE first performs a targeted weight update using the validation set. It then captures the optimization trajectory by calculating the change in neuron activations for each before and after this update. By selecting data whose activation change are most similar to those of the validation set, TRACE ensures alignment in both the representational and optimization domains. Our experiments demonstrate that TRACE outperforms baseline methods across various tasks, particularly in complex, data-scarce scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
相关 Paper
- Capturing the Temporal Dependence of Training Data InfluenceJiachen T. Wang, Dawn Song, James Zou, Prateek Mittal 等ICLR 2025
- Tracing Training Progress: Dynamic Influence Based Selection for Active LearningTianjiao Wan, Kele Xu, Long Lan, Zijian Gao 等ACM MM 2024 · 被引用 3 次
- NICE Data Selection for Instruction Tuning in LLMs with Non-differentiable Evaluation MetricJingtan Wang, Xiaoqiang Lin, Rui Qiao, Pang Wei Koh 等ICML 2025
- Understanding Data Influence in Reinforcement FinetuningHaoru Tan, Xiuzhe Wu, Sitong Wu, Shaofeng Zhang 等NeurIPS 2025 · 被引用 4 次
- Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal ActivationsDa Ma, Gonghu Shang, Zhi Chen, Libo Qin 等NeurIPS 2025 · 被引用 6 次
