Most Influential Subset Selection: Challenges, Promises, and Beyond
Yuzheng Hu, Pingbang Hu, Han Zhao, Jiaqi W. Ma
摘要
How can we attribute the behaviors of machine learning models to their training data? While the classic influence function sheds light on the impact of individual samples, it often fails to capture the more complex and pronounced collective influence of a set of samples. To tackle this challenge, we study the Most Influential Subset Selection (MISS) problem, which aims to identify a subset of training samples with the greatest collective influence. We conduct a comprehensive analysis of the prevailing approaches in MISS, elucidating their strengths and weaknesses. Our findings reveal that influence-based greedy heuristics, a dominant class of algorithms in MISS, can provably fail even in linear regression. We delineate the failure modes, including the errors of influence function and the non-additive structure of the collective influence. Conversely, we demonstrate that an adaptive version of these heuristics which applies them iteratively, can effectively capture the interactions among samples and thus partially address the issues. Experiments on real-world datasets corroborate these theoretical findings and further demonstrate that the merit of adaptivity can extend to more complex scenarios such as classification tasks and non-linear neural networks. We conclude our analysis by emphasizing the inherent trade-off between performance and computational efficiency, questioning the use of additive metrics such as the Linear Datamodeling Score, and offering a range of discussions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Group-Level Data Selection for Efficient PretrainingZichun Yu, Fei Peng, Jie Lei, Arnold Overwijk 等NeurIPS 2025 · 被引用 13 次
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningYuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth 等NeurIPS 2025 · 被引用 13 次
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every IterationShaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu 等ICML 2026 · 被引用 12 次
- GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse ProjectionPingbang Hu, Joseph Melkonian, Weijing Tang, Han Zhao 等NeurIPS 2025 · 被引用 12 次
- Better Training Data Attribution via Better Inverse Hessian-Vector ProductsAndrew Wang, Elisa Nguyen, Runshi Yang, Juhan Bae 等NeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper18
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi 等NeurIPS 2022 · 被引用 185 次
- Scaling Up Influence FunctionsAndrea Schioppa, Polina Zablotskaia, David Vilar, Artem SokolovAAAI 2022 · 被引用 149 次
- Interactive Label Cleaning with Example-based ExplanationsStefano Teso, Andrea Bontempelli, Fausto Giunchiglia, Andrea PasseriniNeurIPS 2021 · 被引用 59 次
相关 Paper
- On Second-Order Group Influence Functions for Black-Box PredictionsSamyadeep Basu, Xuchen You, Soheil FeiziICML 2020 · 被引用 28 次
- Adaptive Greedy versus Non-Adaptive Greedy for Influence MaximizationWei Chen, Binghui Peng, Grant Schoenebeck, Biaoshuai TaoAAAI 2020 · 被引用 27 次
- A Collective Learning Framework to Boost GNN Expressiveness for Node ClassificationMengyue Hang, Jennifer Neville, Bruno RibeiroICML 2021 · 被引用 20 次
- Influence-Based Fair Selection for Sample-Discriminative Backdoor AttackQi Wei, Shuo He, Jiahan Zhang, Lei Feng 等AAAI 2025 · 被引用 1 次
- Finding Most Influential SetsLucas D. Konrad, Nikolas KuschnigICML 2026
