Testing Most Influential Sets
Lucas Darius Konrad, Nikolas Kuschnig
摘要
Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings. While recent work identifies these most influential sets, there is no formal way to tell when maximum influence is excessive rather than expected under natural random sampling variation. We address this gap by developing a principled framework for most influential sets. Focusing on linear least-squares, we derive a convenient exact influence formula and identify the extreme value distributions of maximal influence - the heavy-tailed Fréchet for constant-size sets and heavy-tailed data, and the well-behaved Gumbel for growing sets or light tails. This allows us to conduct rigorous hypothesis tests for excessive influence. We demonstrate through applications across economics, biology, and machine learning benchmarks, resolving contested findings and replacing ad-hoc heuristics with rigorous inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Most Influential Subset Selection: Challenges, Promises, and BeyondYuzheng Hu, Pingbang Hu, Han Zhao, Jiaqi W. MaNeurIPS 2024 · 被引用 39 次
- Theoretical and Practical Perspectives on what Influence Functions DoAndrea Schioppa, Katja Filippova, Ivan Titov, Polina ZablotskaiaNeurIPS 2023 · 被引用 38 次
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin 等NeurIPS 2022 · 被引用 36 次
- "What Data Benefits My Classifier?" Enhancing Model Performance and Interpretability through Influence-Based Data SelectionAnshuman Chhabra, Peizhao Li, Prasant Mohapatra, Hongfu LiuICLR 2024 · 被引用 32 次
- On Second-Order Group Influence Functions for Black-Box PredictionsSamyadeep Basu, Xuchen You, Soheil FeiziICML 2020 · 被引用 28 次
相关 Paper
- Contribution Maximization in Probabilistic DatalogTova Milo, Yuval Moskovitch, Brit YoungmannICDE 2020 · 被引用 2 次
- Why does Throwing Away Data Improve Worst-Group Error?Kamalika Chaudhuri, Kartik Ahuja, Martín Arjovsky, David Lopez-PazICML 2023 · 被引用 27 次
- Robustness Auditing for Linear Regression: To Singularity and BeyondIttai Rubinstein, Samuel B. HopkinsICLR 2025
- Less Is Better: Unweighted Data Subsampling via Influence FunctionZifeng Wang, Hong Zhu, Zhenhua Dong, Xiuqiang He 等AAAI 2020 · 被引用 61 次
- Post-selection inference with HSIC-LassoTobias Freidling, Benjamin Poignard, Héctor Climente-González, Makoto YamadaICML 2021 · 被引用 17 次
