Self-Influence Guided Data Reweighting for Language Model Pre-training
Megh Thakkar, Tolga Bolukbasi, Sriram Ganapathy, Shikhar Vashishth, Sarath Chandar, Partha P. Talukdar
摘要
Language Models (LMs) pre-trained with selfsupervision on large text corpora have become the default starting point for developing models for various NLP tasks. Once the pre-training corpus has been assembled, all data samples in the corpus are treated with equal importance during LM pre-training. However, due to varying levels of relevance and quality of data, equal importance to all the data samples may not be the optimal choice. While data reweighting has been explored in the context of task-specific supervised learning and LM fine-tuning, model-driven reweighting for pretraining data has not been explored. We fill this important gap and propose PRESENCE, a method for jointly reweighting samples by leveraging self-influence (SI) scores as an indicator of sample importance and pre-training. PRESENCE promotes novelty and stability for model pre-training. Through extensive analysis spanning multiple model sizes, datasets, and tasks, we present PRESENCE as an important first step in the research direction of sample reweighting for pre-training language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence ModelsZichun Yu, Spandan Das, Chenyan XiongNeurIPS 2024 · 被引用 117 次
- DOGE: Domain Reweighting with Generalization EstimationSimin Fan, Matteo Pagliardini, Martin JaggiICML 2024 · 被引用 79 次
- How to Train Your LLM Web Agent: A Statistical DiagnosisDheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei 等NeurIPS 2025 · 被引用 19 次
- LayerIF: Estimating Layer Quality for Large Language Models using Influence FunctionsHadi Askari, Shivanshu Gupta, Fei Wang, Anshuman Chhabra 等NeurIPS 2025 · 被引用 16 次
- Enhancing Training Data Attribution with Representational OptimizationWeiwei Sun, Haokun Liu, Nikhil Kandpal, Colin A. Raffel 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper15
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
相关 Paper
- Dynamic Loss-Based Sample Reweighting for Improved Large Language Model PretrainingDaouda Sow, Herbert Woisetschläger, Saikiran Bulusu, Shiqiang Wang 等ICLR 2025
- Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline MethodsWanru Zhao, Yihong Chen, Yuzhi Tang, Wentao Ma 等ICLR 2026 · 被引用 4 次
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model PretrainingJie Hao, Rui Yu, Wei Zhang, Huixia Judy Wang 等ICML 2026 · 被引用 2 次
- Importance Weighting Can Help Large Language Models Self-ImproveChunyang Jiang, Chi-Min Chan, Wei Xue, Qifeng Liu 等AAAI 2025 · 被引用 12 次
- Learning from Noisy Labels via Self-Taught On-the-Fly Meta Loss RescalingMichael Heck, Christian Geishauser, Nurul Lubis, Carel van Niekerk 等AAAI 2025 · 被引用 3 次
