RSDA: Restoring Stale Data Affinity via Dynamic Renovation Strategy for Mitigating Data Scarcity
Yidan Liang, Jia Zhu, Weijie Shi, Hanghui Guo, Yue Cui, Jiawei Shen, Guoqing Ma, Jingjiang Liu, Qingyu Niu, Yilin Wang, Shimin Di, Jiajie Xu
摘要
High-quality data is the cornerstone of advancing large language models. However, the field currently faces a critical dilemma: the supply of premium data is nearing depletion, while vast stale corpora remain underutilized. Our empirical analysis reveals that training models on such data directly often leads to performance degradation. We attribute this phenomenon to the data affinity gap, a misalignment stemming from the model's inability to effectively comprehend the data or inherent quality defects. To bridge this gap, we propose Restoring Stale Data Affinity (RSDA) framework. First, utilizing our proposed potential entropy metric, RSDA quantifies the latent value of samples to effectively identify stale data with higher renovation potential. Subsequently, the framework employs a dynamic renovation strategy selection mechanism to determine the optimal component-level strategy for each instance, transforming low-affinity stale samples into high-quality training data. Comprehensive experimental results demonstrate that RSDA effectively enhances data affinity, achieving performance improvements using less than 10% of the data volume, thereby underscoring that the latent potential of stale corpora remains largely untapped. Our code is available at https://github.com/wenfiii/RSDA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-FoldAmrith Setlur, Saurabh Garg, Xinyang Geng, Naman Garg 等NeurIPS 2024 · 被引用 143 次
相关 Paper
- LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMsJianghao Chen, Junhong Wu, Yangyifan Xu, Jiajun ZhangACL 2025
- Rethinking the Reversal Curse of LLMs: a Prescription from Human Knowledge ReversalZhicong Lu, Li Jin, Peiguang Li, Yu Tian 等EMNLP 2024 · 被引用 1 次
- DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language ModelsRanchi Zhao, Zhen Leng Thai, Yifan Zhang, Shengding Hu 等EMNLP 2024
- Data Rejuvenation: Exploiting Inactive Training Examples for Neural Machine TranslationWenxiang Jiao, Xing Wang, Shilin He, Irwin King 等EMNLP 2020 · 被引用 18 次
- RePro: Training Language Models to Faithfully Recycle the Web for PretrainingZichun Yu, Chenyan XiongICML 2026 · 被引用 2 次
