Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and How
Sebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter, Josif Grabocka
Abstract
With the ever-increasing number of pretrained models, machine learning practitioners are continuously faced with the decision of which pretrained model to use, and how to finetune it for a new dataset. In this paper, we propose a methodology that jointly searches for the optimal pretrained model and the hyperparameters for finetuning it. Our method transfers knowledge about the performance of many pretrained models with multiple hyperparameter configurations on a series of datasets. To this aim, we evaluated over 20k hyperparameter configurations for finetuning 24 pretrained image classification models on 87 datasets to generate a large-scale meta-dataset. We meta-learn a gray-box performance predictor on the learning curves of this meta-dataset and use it for fast hyperparameter optimization on new datasets. We empirically demonstrate that our resulting approach can quickly select an accurate pretrained model for a new dataset together with its optimal hyperparameters. To facilitate reproducibility, we open-source our code and release our meta-dataset. 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 790cd5a1-bb89-4fbf-87ec-893e1998b6d1Cited by top-tier papers7
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter TuningDong Bok Lee, Aoxuan Silvia Zhang, Byungjoo Kim, Junhyeon Park et al.NeurIPS 2025 · 2 citations
- Implicit Modeling for Transferability Estimation of Vision Foundation ModelsYaoyan Zheng, Huiqun Wang, Nan Zhou, Di HuangNeurIPS 2025 · 1 citation
- Varying Shades of Wrong: Aligning LLMs with Wrong Answers OnlyJihan Yao, Wenxuan Ding, Shangbin Feng, Lucy Lu Wang et al.ICLR 2025
- OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?Liangze Jiang, Damien TeneyICML 2025
- Efficient Cross-Episode Meta-RLGresa Shala, André Biedenkapp, Pierre Krack, Florian Walter et al.ICLR 2025
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- Zero-shot AutoML with Pretrained ModelsEkrem Öztürk, Fabio Ferreira, Hadi S. Jomaa, Lars Schmidt-Thieme et al.ICML 2022 · 17 citations
- Rapid Neural Architecture Search by Learning to Generate Graphs from DatasetsHayeon Lee, Eunyoung Hyung, Sung Ju HwangICLR 2021 · 57 citations
- Landmark-Guided Policy Optimization for Multi-Objective Language Model SelectionMarcio Monteiro, Weichen Li, Puyu Wang, Marius Kloft et al.ICML 2026
- Model Selection with Model Zoo via Graph LearningZiyu Li, Hilco van der Wilk, Danning Zhan, Megha Khosla et al.ICDE 2024 · 6 citations
- Co-Tuning for Transfer LearningKaichao You, Zhi Kou, Mingsheng Long, Jianmin WangNeurIPS 2020 · 105 citations
