MTAdam: Automatic Balancing of Multiple Training Loss Terms
Itzik Malkiel, Lior Wolf
Abstract
When training neural models, it is common to combine multiple loss terms. The balancing of these terms requires considerable human effort and is computationally demanding. Moreover, the optimal trade-off between the loss terms can change as training progresses, e.g., for adversarial terms. In this work, we generalize the Adam optimization algorithm to handle multiple loss terms. The guiding principle is that for every layer, the gradient magnitude of the terms should be balanced. To this end, the Multi-Term Adam (MTAdam) computes the derivative of each loss term separately, infers the first and second moments per parameter and loss term, and calculates a first moment for the magnitude per layer of the gradients arising from each loss. This magnitude is used to continuously balance the gradients across all layers, in a manner that both varies from one layer to the next and dynamically changes over time. Our results show that training with the new method leads to fast recovery from suboptimal initial loss weighting and to training outcomes that match or improve conventional training with the prescribed hyperparameters of each method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48d74ea7-a701-476a-8fe2-e794f2a1b399Cited by top-tier papers3
- Multi-behavior Self-supervised Learning for RecommendationJingcao Xu, Chaokun Wang, Cheng Wu, Yang Song et al.SIGIR 2023 · 80 citations
- MetaBalance: Improving Multi-Task Recommendations via Adapting Gradient Magnitudes of Auxiliary TasksYun He, Xue Feng, Cheng Cheng, Geng Ji et al.WWW 2022 · 69 citations
- Rep-MTL: Unleashing the Power of Representation-Level Task Saliency for Multi-Task LearningZedong Wang, Siyuan Li, Dan XuICCV 2025 · 3 citations
Related papers
- IO-Adam: Rethinking Memory-Efficient Adaptive Optimizers from Gradient ComputationYiting Chen, Zongwei Huo, Junchi YanICML 2026
- AdaTask: A Task-Aware Adaptive Learning Rate Approach to Multi-Task LearningEnneng Yang, Junwei Pan, Ximei Wang, Haibin Yu et al.AAAI 2023 · 70 citations
- A Method for Enhancing Generalization of Adam by Multiple IntegrationsLong Jin, Han Nong, Liangming Chen, Zhenming SuAAAI 2025 · 2 citations
- The AdEMAMix Optimizer: Better, Faster, OlderMatteo Pagliardini, Pierre Ablin, David GrangierICLR 2025
- Bayesian filtering unifies adaptive and non-adaptive neural network optimization methodsLaurence AitchisonNeurIPS 2020 · 23 citations
