Lune

ICML2026Top-tier venue

Clustered Influence Functions

Miklós Máté Badó, Kristian Fenech

2026Year

Abstract

Influence functions are a standard tool for data debugging and unlearning, but they become impractical for high-query subset workloads such as large-KK cross-validation, repeated resampling, or interactive what-if analysis as each subset query typically requires an expensive inverse-curvature solve. We introduce Clustered Influence Functions (CiF), which turns subset influence into an amortized subset oracle. We build a compact cache once by clustering training gradients, solve a damped Generalised Gauss-Newton system only for cluster means, and answer new subset queries by a linear recombination using cluster membership counts. This yields per-query cost of O(Cp)O(Cp) linear in the cache size CC, and the number of model parameters pp. We further provide a diagnostic error bound that decomposes approximation error into a clustering scatter term and a solver residual term, making the accuracy-compute tradeoff explicit through the cache budget and solver tolerance. Evaluations across MNIST, CIFAR-10 show that CiF matches per-query influence rankings while significantly reducing the total runtime in high-QQ regimes, enabling influence-based workflows that are otherwise computationally prohibitive.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines