What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff G. Schneider, Eduard H. Hovy, Roger Baker Grosse, Eric P. Xing
Abstract
Large language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution 2 ), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation methods to recent LLMs and their vast training datasets has been largely limited by prohibitive compute and memory costs. In this work, we focus on influence functions, a popular gradient-based data valuation method, and significantly improve its scalability with an efficient gradient projection strategy called LOGRA that leverages the gradient structure in backpropagation. We then provide a theoretical motivation of gradient projection approaches to influence functions to promote trust in the data valuation process. Lastly, we lower the barrier to implementing data valuation systems by introducing LOGIX, a software package that can transform existing training code into data valuation code with minimal effort. In our data valuation experiments, LOGRA achieves competitive accuracy against more expensive baselines while showing up to 6,500× improvement in throughput and 5× reduction in GPU memory usage when applied to Llama3-8B-Instruct and the 1B-token dataset (open source project: link).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee3e41de-902a-437c-8bc2-27a1ce1d9688Cited by top-tier papers41
- GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K ResolutionFengxiang Wang, Mingshuo Chen, Yueying Li, Di Wang et al.NeurIPS 2025 · 46 citations
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 41 citations
- LayerIF: Estimating Layer Quality for Large Language Models using Influence FunctionsHadi Askari, Shivanshu Gupta, Fei Wang, Anshuman Chhabra et al.NeurIPS 2025 · 16 citations
- LEAD: Iterative Data Selection for Efficient LLM Instruction TuningXiaotian Lin, Yanlin Qi, Yizhang Zhu, Themis Palpanas et al.VLDB 2026 · 16 citations
- Bayesian Influence Functions for Hessian-Free Data AttributionPhilipp Alexander Kreer, Wilson Wu, Maxwell Adam, Zach Furman et al.ICLR 2026 · 14 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
Related papers
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 112 citations
- For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMsWenlong Deng, Qi Zeng, Jiaming Zhang, Minghui Chen et al.ACL 2026
- Enhancing Training Data Attribution with Representational OptimizationWeiwei Sun, Haokun Liu, Nikhil Kandpal, Colin A. Raffel et al.NeurIPS 2025 · 9 citations
- Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuningSirui Chen, Yunzhe Qi, Mengting Ai, Yifan Sun et al.ICLR 2026 · 9 citations
- RRInf: Efficient Influence Function Estimation via Ridge Regression for Large Language Models and Text-to-Image Diffusion ModelsZhuozhuo Tu, Cheng Chen, Yuxuan DuEMNLP 2025
