Lune

ICML2026顶会

Clustered Influence Functions

Miklós Máté Badó, Kristian Fenech

2026年份

摘要

Influence functions are a standard tool for data debugging and unlearning, but they become impractical for high-query subset workloads such as large-KK cross-validation, repeated resampling, or interactive what-if analysis as each subset query typically requires an expensive inverse-curvature solve. We introduce Clustered Influence Functions (CiF), which turns subset influence into an amortized subset oracle. We build a compact cache once by clustering training gradients, solve a damped Generalised Gauss-Newton system only for cluster means, and answer new subset queries by a linear recombination using cluster membership counts. This yields per-query cost of O(Cp)O(Cp) linear in the cache size CC, and the number of model parameters pp. We further provide a diagnostic error bound that decomposes approximation error into a clustering scatter term and a solver residual term, making the accuracy-compute tradeoff explicit through the cache budget and solver tolerance. Evaluations across MNIST, CIFAR-10 show that CiF matches per-query influence rankings while significantly reducing the total runtime in high-QQ regimes, enabling influence-based workflows that are otherwise computationally prohibitive.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖