For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
Wenlong Deng, Qi Zeng, Jiaming Zhang, Minghui Chen, Zixin Ding, Christos Thrampoulidis, Boying Gong, Xiaoxiao Li
Abstract
Data valuation is essential for enhancing the transparency and accountability of large language models (LLMs) and vision-language models (VLMs). However, existing methods typically rely on gradient computations, making them computationally prohibitive for billion-parameter models and precluding batch parallelization. In this work, we introduce For-Value, a forward-only data valuation framework that enables efficient batch-scalable value estimation while maintaining effectiveness. Leveraging the expressive power of pretrained LLMs/VLMs, we theoretically demonstrate that data valuation can be captured by the alignment between the final hidden representations and prediction errors at the last layer. In light of this insight, For-Value computes data value using a simple closed-form expression with a single forward pass, eliminating the need for costly backpropagation and enabling efficient batch calculating at scale. Extensive experiments show that For-Value matches or outperforms gradient-based baselines in detecting influential data and mislabeled data, while achieving significant efficiency improvements. Our code is available at GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e3953ad-6aaf-4f2b-9070-ad29cdc3ab38Builds on13
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 112 citations
- Resolving Training Biases via Influence-based Data RelabelingShuming Kong, Yanyan Shen, Linpeng HuangICLR 2022 · 71 citations
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 41 citations
Related papers
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao et al.NeurIPS 2025 · 112 citations
- Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James ZouICML 2023 · 54 citations
- The Mirrored Influence Hypothesis: Efficient Data Influence Estimation by Harnessing Forward PassesMyeongseob Ko, Feiyang Kang, Weiyan Shi, Ming Jin et al.CVPR 2024
- EcoVal: An Efficient Data Valuation Framework for Machine LearningAyush K. Tarun, Vikram S. Chundawat, Murari Mandal, Hong Ming Tan et al.KDD 2024 · 3 citations
- Understanding Data Influence in Reinforcement FinetuningHaoru Tan, Xiuzhe Wu, Sitong Wu, Shaofeng Zhang et al.NeurIPS 2025 · 4 citations
