PROFIT: A Specialized Optimizer for Deep Fine Tuning
Anirudh Srinivasan Chakravarthy, Shuai Kyle Zheng, Xin Huang, Sachithra Hemachandra, Xiao Zhang, Yuning Chai, Zhao Chen
Abstract
The fine-tuning of pre-trained models has become ubiquitous in generative AI, computer vision, and robotics. Although much attention has been paid to improving the efficiency of fine-tuning models, there has been less scholarship around fine-tuning specifically for improved model performance. To remedy this gap, we present PROFIT, one of the first optimizers designed to incrementally fine-tune converged models on new tasks and/or datasets. Unlike traditional optimizers such as SGD or Adam, which make minimal assumptions due to random initializations, PROFIT takes the properties of a converged model into account explicitly to regularize the optimization process. Employing a temporal gradient-orthogonalization process, PROFIT outperforms fine-tuning methods in various tasks, from image classification to multimodal language model training to large-scale motion prediction. Moreover, PROFIT is encapsulated as a modular optimizer, which makes it easy to integrate directly into any training pipeline with minimal engineering effort. maintaining performance on old tasks. We will also later show how to overcome this constraint even in non-proximal settings by introducing a warmup phase.
PROFIT (PROximal FIne Tuning) is shown schematically in Fig. 1. To the best of our knowledge, PROFIT is among the first optimizers explicitly designed for fine-tuning.
Our main contributions are as follows.
• We introduce PROFIT, an optimizer for fine-tuning converged models, easily integrated into any deep learning framework.
• We show that PROFIT allows unsupervised training as if the original data were available.
• We show that PROFIT outperforms standard fine-tuning methods on various tasks, from image classification to VLM fine-tuning to large-scale motion prediction for autonomous driving.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Two-stage LLM Fine-tuning with Less Specialization and More GeneralizationYihan Wang, Si Si, Daliang Li, Michal Lukasik et al.ICLR 2024 · 45 citations
- MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic OptimizationDa Chang, Ganzhao YuanNeurIPS 2025 · 9 citations
- Fast Trainable Projection for Robust Fine-tuningJunjiao Tian, Yen-Cheng Liu, James Seale Smith, Zsolt KiraNeurIPS 2023 · 23 citations
- Multi-Token Prediction Needs RegistersAnastasios Gerontopoulos, Spyridon Gidaris, Nikos KomodakisNeurIPS 2025 · 13 citations
- Towards Efficient Low-Order Hybrid Optimizer for Language Model Fine-TuningMinping Chen, You-Liang Huang, Zeyi WenAAAI 2025 · 6 citations
