You Only Train Once: Loss-Conditional Training of Deep Networks
Alexey Dosovitskiy, Josip Djolonga
Abstract
In many machine learning problems, loss functions are weighted sums of several terms. A typical approach to dealing with these is to train multiple separate models with different selections of weights and then either choose the best one according to some criterion or keep multiple models if it is desirable to maintain a diverse set of solutions. This is inefficient both at training and at inference time. We propose a method that allows replacing multiple models trained on one loss function each by a single model trained on a distribution of losses. At test time a model trained this way can be conditioned to generate outputs corresponding to any loss from the training distribution of losses. We demonstrate this approach on three tasks with parametrized losses: beta-VAE, learned image compression, and fast style transfer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 263 citations
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 189 citations
- Pareto Set Learning for Expensive Multi-Objective OptimizationXi Lin, Zhiyuan Yang, Xiaoyuan Zhang, Qingfu ZhangNeurIPS 2022 · 119 citations
- Multi-Instance Pose Networks: Rethinking Top-Down Pose EstimationRawal Khirodkar, Visesh Chari, Amit Agrawal, Ambrish TyagiICCV 2021 · 80 citations
- Efficient Split-Mix Federated Learning for On-Demand and In-Situ CustomizationJunyuan Hong, Haotao Wang, Zhangyang Wang, Jiayu ZhouICLR 2022 · 77 citations
Builds on1
Related papers
- Tunable Convolutions with Parametric Multi-Loss OptimizationMatteo Maggioni, Thomas Tanay, Francesca Babiloni, Steven McDonagh et al.CVPR 2023
- Learning useful representations for shifting tasks and distributionsJianyu Zhang, Léon BottouICML 2023 · 21 citations
- Multi-Rate VAE: Train Once, Get the Full Rate-Distortion CurveJuhan Bae, Michael R. Zhang, Michael Ruan, Eric Wang et al.ICLR 2023 · 3 citations
- Soup-of-Experts: Pretraining Specialist Models via Parameters AveragingPierre Ablin, Angelos Katharopoulos, Skyler Seto, David GrangierICML 2025
- Sym-Parameterized Dynamic Inference for Mixed-Domain Image TranslationSimyung Chang, Seonguk Park, John Yang, Nojun KwakICCV 2019 · 8 citations
