Sketchy Moment Matching: Toward Fast and Provable Data Selection for Finetuning
Yijun Dong, Viet Hoang Phan, Xiang Pan, Qi Lei
Abstract
We revisit data selection in a modern context of finetuning from a fundamental perspective. Extending the classical wisdom of variance minimization in low dimensions to high-dimensional finetuning, our generalization analysis unveils the importance of additionally reducing bias induced by low-rank approximation. Inspired by the variance-bias tradeoff in high dimensions from the theory, we introduce Sketchy Moment Matching (SkMM), a scalable data selection scheme with two stages. (i) First, the bias is controlled using gradient sketching that explores the finetuning parameter space for an informative low-dimensional subspace S; (ii) then the variance is reduced over S via moment matching between the original and selected datasets. Theoretically, we show that gradient sketching is fast and provably accurate: selecting n samples by reducing variance over S preserves the fast-rate generalization O(dim(S)/n), independent of the parameter dimension. Empirically, we concretize the variance-bias balance via synthetic experiments and demonstrate the effectiveness of SkMM for finetuning in real vision tasks. * Equal contribution. 2 Throughout this work, we refer to "low-dimension" as the setting where the number of finetuning parameters r is smaller than the selected downstream sample size n, while "high-dimension" refers to the opposite, r > n.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2298c40-abb5-4232-8e20-498bbae6b754Cited by top-tier papers3
- LAMDAS: LLM as an Implicit Classifier for Domain-specific Data SelectionJian Wu, Hang Yu, Bingchang Liu, Wenjie Yang et al.AAAI 2026 · 1 citation
- Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic DimensionYijun Dong, Yicheng Li, Yunai Li, Jason D. Lee et al.ICML 2025
- PEAKS: Selecting Key Training Examples Incrementally via Prediction Error Anchored by Kernel SimilarityMustafa Burak Gurbuz, Xingyu Zheng, Constantine DovrolisICML 2025
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
Related papers
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian et al.NeurIPS 2022 · 69 citations
- Sketched Adaptive Distributed Deep Learning: A Sharp Convergence AnalysisZhijie Chen, Qiaobo Li, Arindam BanerjeeNeurIPS 2025 · 1 citation
- A Fast and Accurate Estimator for Large Scale Linear Model via Data AveragingRui Wang, Yanyan Ouyang, Panpan Yu, Wangli XuNeurIPS 2023 · 1 citation
- Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language ModelsLijing Wang, Yingya Li, Timothy Miller, Steven Bethard et al.ACL 2023 · 10 citations
- Low-Rank Rescaled Vision Transformer Fine-Tuning: A Residual Design ApproachWei Dong, Xing Zhang, Bihui Chen, Dawei Yan et al.CVPR 2024
