A Medical Data-Effective Learning Benchmark for Highly Efficient Pre-training of Foundation Models
Wenxuan Yang, Weimin Tan, Yuqi Sun, Bo Yan
摘要
Foundation models, pre-trained on massive datasets, have achieved unprecedented generalizability. However, is it truly necessary to involve such vast amounts of data in pre-training, consuming extensive computational resources? This paper introduces data-effective learning, aiming to use data in the most impactful way to pre-train foundation models. This involves strategies that focus on data quality rather than quantity, ensuring the data used for training has high informational value. Data-effective learning plays a profound role in accelerating foundation model training, reducing computational costs, and saving data storage, which is very important as the volume of medical data in recent years has grown beyond many people's expectations. However, due to the lack of standards and comprehensive benchmark, research on medical data-effective learning is poorly studied. To address this gap, our paper introduces a comprehensive benchmark specifically for evaluating data-effective learning in the medical field. This benchmark includes a dataset with millions of data samples from 31 medical centers (DataDEL), a baseline method for comparison (MedDEL), and a new evaluation metric (NormDEL) to objectively measure data-effective learning performance. Our extensive experimental results show the baseline MedDEL can achieve performance comparable to the original large dataset with only 5% of the data. Establishing such an open data-effective learning benchmark is crucial for the medical foundation model research community because it facilitates efficient data use, promotes collaborative breakthroughs, and fosters the development of cost-effective, scalable, and impactful healthcare solutions. The benchmark can be accessed at https://github.com/shadow2469/Data-Effective-Learning-A-Comprehensive-Medical-Benchmark.git GitHub Repository.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung 等ACL 2026 · 被引用 9 次
- Scaling Laws for Data-Efficient Visual Transfer LearningWenxuan Yang, Qingqv Wei, Chenxi Ma, Weimin Tan 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang 等ICCV 2019 · 被引用 469 次
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li 等CVPR 2022
相关 Paper
- DeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and RoutingConglong Li, Zhewei Yao, Xiaoxia Wu, Minjia Zhang 等AAAI 2024 · 被引用 43 次
- DataRater: Meta-Learned Dataset CurationDan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf 等NeurIPS 2025 · 被引用 17 次
- Predictive Data Selection: The Data That Predicts Is the Data That TeachesKaShun Shum, Yuzhen Huang, Hongjian Zou, Qi Ding 等ICML 2025
- Data Shapley in One Training RunJiachen T. Wang, Prateek Mittal, Dawn Song, Ruoxi JiaICLR 2025
- NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient FrameworkXingcheng Yao, Yanan Zheng, Xiaocong Yang, Zhilin YangICML 2022 · 被引用 50 次
