A Meta-Learning Approach to Predicting Performance and Data Requirements
Achin Jain, Gurumurthy Swaminathan, Paolo Favaro, Hao Yang, Avinash Ravichandran, Hrayr Harutyunyan, Alessandro Achille, Onkar Dabeer, Bernt Schiele, Ashwin Swaminathan, Stefano Soatto
摘要
We propose an approach to estimate the number of samples required for a model to reach a target performance. We find that the power law, the de facto principle to estimate model performance, leads to large error when using a small dataset (e.g., 5 samples per class) for extrapolation. This is because the log-performance error against the log-dataset size follows a nonlinear progression in the few-shot regime followed by a linear progression in the high-shot regime. We introduce a novel piecewise power law (PPL) that handles the two data regimes differently. To estimate the parameters of the PPL, we introduce a random forest regressor trained via meta learning that generalizes across classification/detection tasks, ResNet/ViT based architectures, and random/pre-trained initializations. The PPL improves the performance estimation on average by 37% across 16 classification and 33% across 10 detection datasets, compared to the power law. We further extend the PPL to provide a confidence bound and use it to limit the prediction horizon that reduces over-estimation of data by 76% on classification and 91% on detection datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Scaling Laws for Hyperparameter OptimizationArlind Kadra, Maciej Janowski, Martin Wistuba, Josif GrabockaNeurIPS 2023 · 被引用 23 次
- How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and BudgetGuy Hacohen, Daphna WeinshallNeurIPS 2023 · 被引用 23 次
- Energy-based Automated Model EvaluationRu Peng, Heming Zou, Haobo Wang, Yawen Zeng 等ICLR 2024 · 被引用 18 次
- Scaling Laws for Downstream Task Performance in Machine TranslationBerivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas 等ICLR 2025
- Monotonic Variational Gaussian Process for Efficient Data CollectionDonghyun Lee, Young Myoung KoICML 2026
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 被引用 171 次
- Exploring the Limits of Large Scale Pre-trainingSamira Abnar, Mostafa Dehghani, Behnam Neyshabur, Hanie SedghiICLR 2022 · 被引用 135 次
- Benchopt: Reproducible, efficient and collaborative optimization benchmarksThomas Moreau, Mathurin Massias, Alexandre Gramfort, Pierre Ablin 等NeurIPS 2022 · 被引用 58 次
相关 Paper
- How Much More Data Do I Need? Estimating Requirements for Downstream TasksRafid Mahmood, James Lucas, David Acuna, Daiqing Li 等CVPR 2022 · 被引用 21 次
- A Theoretical Analysis of the Number of Shots in Few-Shot LearningTianshi Cao, Marc T. Law, Sanja FidlerICLR 2020 · 被引用 75 次
- Meta-Learning to Detect Rare ObjectsYu-Xiong Wang, Deva Ramanan, Martial HebertICCV 2019 · 被引用 339 次
- Meta-RCNN: Meta Learning for Few-Shot Object DetectionXiongwei Wu, Doyen Sahoo, Steven C. H. HoiACM MM 2020 · 被引用 94 次
- Model Performance Scaling with Multiple Data SourcesTatsunori HashimotoICML 2021 · 被引用 38 次
