Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuning
Sirui Chen, Yunzhe Qi, Mengting Ai, Yifan Sun, Ruizhong Qiu, Jiaru Zou, Jingrui He
Abstract
Supervised fine-tuning (SFT) relies critically on selecting training data that most benefits model's downstream performance. Gradient-based data selection methods such as TracIn and Influence Functions leverage influence to identify useful samples, but their computational cost scales poorly, making them impractical for multi-billion-parameter large language models (LLMs). A common alternative is to use off-the-shelf smaller models as proxies, but they remain suboptimal since their learning dynamics are unclear, their sizes cannot be flexibly adjusted, and they cannot be further aligned with the target model in terms of gradient-based influence estimation. To address these challenges, we introduce IPROX, a two-stage framework that derives influence-preserving proxies directly from the target model. It first applies a low-rank compression stage to preserve influence information of the target model, and then an aligning stage to align both model gradients and logits, thereby constructing proxies that flexibly control computational cost while retaining the target model's influence. Experimental results across diverse LLM families and evaluation tasks show that IPROX consistently outperforms off-theshelf proxies and baseline methods. On Qwen3-4B, a 1.5B proxy constructed with IPROX achieves stronger performance than the larger 1.7B off-the-shelf proxy. Notably, on Llama3.2, IPROX achieves better performance than baselines while reducing computational cost by more than half relative to the full 3B model. These results show that IPROX provides effective influence-preserving proxies, making gradient-based data selection more scalable for LLMs. The code is available at https://github.com/csr16/IProX * Equal Contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3800c535-c494-4dab-a00a-cda91fccb848Cited by top-tier papers4
- PLANETALIGN: A Comprehensive Python Library for Benchmarking Network AlignmentQi Yu, Zhichen Zeng, Yuchen Yan, Zhining Liu et al.ICLR 2026 · 12 citations
- GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization GeometryGuanghui Min, Tianhao Huang, Ke Wan, Chen ChenICML 2026 · 3 citations
- Graph homophily booster: Reimagining the role of discrete features in heterophilic graph learningRuizhong Qiu, Ting-Wei Li, Gaotang Li, Hanghang TongICLR 2026 · 2 citations
- AvAtar: Learning to Align via Active Optimal TransportQi Yu, Ruizhong Qiu, Zhichen Zeng, My T. Thai et al.ICML 2026 · 1 citation
Builds on29
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
Related papers
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model PretrainingJie Hao, Rui Yu, Wei Zhang, Huixia Judy Wang et al.ICML 2026 · 2 citations
- BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMsChaoyuan Shen, Chi Zhang, Chengliang Chai, Jiacheng Wang et al.VLDB 2026 · 2 citations
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 112 citations
- FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware FusionTao Fan, Guoqiang Ma, Yuanfeng Song, Lixin Fan et al.ACL 2026
- Efficient Data Selection at Scale via Influence DistillationMahdi Nikdan, Vincent Cohen-Addad, Dan Alistarh, Vahab MirrokniNeurIPS 2025 · 15 citations
