Efficient Estimation of Kernel Surrogate Models for Task Attribution
Zhenshuo Zhang, Minxuan Duan, Hongyang R. Zhang
摘要
Modern AI agents such as large language models are trained on diverse tasks---translation, code generation, mathematical reasoning, and text prediction---simultaneously. A key question is how to quantify the influence of each individual training task on performance on a target task, a problem we refer to as task attribution. The direct approach, leave-one-out retraining, measures the effect of removing each task, but is computationally infeasible at scale. An alternative approach that builds surrogate models to predict the performance on a target task for any subset of training tasks has emerged in the recent literature. Prior work focuses on linear surrogate models, which capture first-order relationships but miss nonlinear interactions such as synergy, antagonism, or XOR-type effects. In this paper, we first consider a unified task-weighting framework for analyzing task-attribution methods and establish a new connection between linear surrogate models and influence functions via a second-order analysis. Then, we introduce kernel surrogate models, which more effectively represent second-order task interactions. To efficiently learn the kernel surrogate, we develop a gradient-based estimation procedure that leverages a first-order approximation of pretrained models; empirically, this yields accurate surrogate estimates with less than % relative error without repeated retraining. Experiments across multiple domains---including mathematical reasoning in transformers, in-context learning, and multi-objective reinforcement learning---demonstrate the effectiveness of kernel surrogate models. They achieve a % higher correlation with the leave-one-out ground truth than linear surrogates and influence-function baselines, enabling more accurate and scalable task attribution. When used for downstream task selection, kernel surrogate models further yield a % improvement in demonstration selection for in-context learning and multi-objective reinforcement learning benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMsYuxuan Li, Aoi Naito, Hirokazu ShiradoICML 2026 · 被引用 9 次
- WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle PointsDongyue Li, Zechun Liu, Kai Yi, Zhenshuo Zhang 等ICML 2026
它引用的顶会 Paper19
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
- Understanding and Improving Information Transfer in Multi-Task LearningSen Wu, Hongyang R. Zhang, Christopher RéICLR 2020 · 被引用 183 次
相关 Paper
- DETAIL: Task DEmonsTration Attribution for Interpretable In-context LearningZijian Zhou, Xiaoqiang Lin, Xinyi Xu, Alok Prakash 等NeurIPS 2024 · 被引用 9 次
- AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context AttributionFengyuan Liu, Nikhil Kandpal, Colin RaffelICLR 2025
- Enhancing Training Data Attribution with Representational OptimizationWeiwei Sun, Haokun Liu, Nikhil Kandpal, Colin A. Raffel 等NeurIPS 2025 · 被引用 9 次
- Linear-Time Demonstration Selection for In-Context Learning via Gradient EstimationZiniu Zhang, Zhenshuo Zhang, Dongyue Li, Lu Wang 等EMNLP 2025
- Scalable Influence and Fact Tracing for Large Language Model PretrainingTyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon 等ICLR 2025 · 被引用 1 次
