Most Influential Subset Selection: Challenges, Promises, and Beyond
Yuzheng Hu, Pingbang Hu, Han Zhao, Jiaqi W. Ma
Abstract
How can we attribute the behaviors of machine learning models to their training data? While the classic influence function sheds light on the impact of individual samples, it often fails to capture the more complex and pronounced collective influence of a set of samples. To tackle this challenge, we study the Most Influential Subset Selection (MISS) problem, which aims to identify a subset of training samples with the greatest collective influence. We conduct a comprehensive analysis of the prevailing approaches in MISS, elucidating their strengths and weaknesses. Our findings reveal that influence-based greedy heuristics, a dominant class of algorithms in MISS, can provably fail even in linear regression. We delineate the failure modes, including the errors of influence function and the non-additive structure of the collective influence. Conversely, we demonstrate that an adaptive version of these heuristics which applies them iteratively, can effectively capture the interactions among samples and thus partially address the issues. Experiments on real-world datasets corroborate these theoretical findings and further demonstrate that the merit of adaptivity can extend to more complex scenarios such as classification tasks and non-linear neural networks. We conclude our analysis by emphasizing the inherent trade-off between performance and computational efficiency, questioning the use of additive metrics such as the Linear Datamodeling Score, and offering a range of discussions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Group-Level Data Selection for Efficient PretrainingZichun Yu, Fei Peng, Jie Lei, Arnold Overwijk et al.NeurIPS 2025 · 13 citations
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningYuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth et al.NeurIPS 2025 · 13 citations
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every IterationShaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu et al.ICML 2026 · 12 citations
- GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse ProjectionPingbang Hu, Joseph Melkonian, Weijing Tang, Han Zhao et al.NeurIPS 2025 · 12 citations
- Better Training Data Attribution via Better Inverse Hessian-Vector ProductsAndrew Wang, Elisa Nguyen, Runshi Yang, Juhan Bae et al.NeurIPS 2025 · 12 citations
Builds on18
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc et al.ICML 2023 · 260 citations
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
- Scaling Up Influence FunctionsAndrea Schioppa, Polina Zablotskaia, David Vilar, Artem SokolovAAAI 2022 · 149 citations
- Interactive Label Cleaning with Example-based ExplanationsStefano Teso, Andrea Bontempelli, Fausto Giunchiglia, Andrea PasseriniNeurIPS 2021 · 59 citations
Related papers
- On Second-Order Group Influence Functions for Black-Box PredictionsSamyadeep Basu, Xuchen You, Soheil FeiziICML 2020 · 28 citations
- Adaptive Greedy versus Non-Adaptive Greedy for Influence MaximizationWei Chen, Binghui Peng, Grant Schoenebeck, Biaoshuai TaoAAAI 2020 · 27 citations
- A Collective Learning Framework to Boost GNN Expressiveness for Node ClassificationMengyue Hang, Jennifer Neville, Bruno RibeiroICML 2021 · 20 citations
- Influence-Based Fair Selection for Sample-Discriminative Backdoor AttackQi Wei, Shuo He, Jiahan Zhang, Lei Feng et al.AAAI 2025 · 1 citation
- Finding Most Influential SetsLucas D. Konrad, Nikolas KuschnigICML 2026
