You Only Train Once: Loss-Conditional Training of Deep Networks
Alexey Dosovitskiy, Josip Djolonga
摘要
In many machine learning problems, loss functions are weighted sums of several terms. A typical approach to dealing with these is to train multiple separate models with different selections of weights and then either choose the best one according to some criterion or keep multiple models if it is desirable to maintain a diverse set of solutions. This is inefficient both at training and at inference time. We propose a method that allows replacing multiple models trained on one loss function each by a single model trained on a distribution of losses. At test time a model trained this way can be conditioned to generate outputs corresponding to any loss from the training distribution of losses. We demonstrate this approach on three tasks with parametrized losses: beta-VAE, learned image compression, and fast style transfer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 被引用 263 次
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 被引用 189 次
- Pareto Set Learning for Expensive Multi-Objective OptimizationXi Lin, Zhiyuan Yang, Xiaoyuan Zhang, Qingfu ZhangNeurIPS 2022 · 被引用 119 次
- Multi-Instance Pose Networks: Rethinking Top-Down Pose EstimationRawal Khirodkar, Visesh Chari, Amit Agrawal, Ambrish TyagiICCV 2021 · 被引用 80 次
- Efficient Split-Mix Federated Learning for On-Demand and In-Situ CustomizationJunyuan Hong, Haotao Wang, Zhangyang Wang, Jiayu ZhouICLR 2022 · 被引用 77 次
它引用的顶会 Paper1
相关 Paper
- Tunable Convolutions with Parametric Multi-Loss OptimizationMatteo Maggioni, Thomas Tanay, Francesca Babiloni, Steven McDonagh 等CVPR 2023
- Learning useful representations for shifting tasks and distributionsJianyu Zhang, Léon BottouICML 2023 · 被引用 21 次
- Multi-Rate VAE: Train Once, Get the Full Rate-Distortion CurveJuhan Bae, Michael R. Zhang, Michael Ruan, Eric Wang 等ICLR 2023 · 被引用 3 次
- Soup-of-Experts: Pretraining Specialist Models via Parameters AveragingPierre Ablin, Angelos Katharopoulos, Skyler Seto, David GrangierICML 2025
- Sym-Parameterized Dynamic Inference for Mixed-Domain Image TranslationSimyung Chang, Seonguk Park, John Yang, Nojun KwakICCV 2019 · 被引用 8 次
