Bayesian Uncertainty for Gradient Aggregation in Multi-Task Learning
Idan Achituve, Idit Diamant, Arnon Netzer, Gal Chechik, Ethan Fetaya
摘要
As machine learning becomes more prominent there is a growing demand to perform several inference tasks in parallel. Running a dedicated model for each task is computationally expensive and therefore there is a great interest in multi-task learning (MTL). MTL aims at learning a single model that solves several tasks efficiently. Optimizing MTL models is often achieved by first computing a single gradient per task and then aggregating the gradients for obtaining a combined update direction. However, this approach do not consider an important aspect, the sensitivity in the gradient dimensions. Here, we introduce a novel gradient aggregation approach using Bayesian inference. We place a probability distribution over the task-specific parameters, which in turn induce a distribution over the gradients of the tasks. This additional valuable information allows us to quantify the uncertainty in each of the gradients dimensions, which can then be factored in when aggregating them. We empirically demonstrate the benefits of our approach in a variety of datasets, achieving state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task DifficultyYanqi Dai, Yong Wang, Zebin You, Dong Jing 等WWW 2026 · 被引用 4 次
- LexTempus: Enhancing Temporal Generalizability of Legal Language Models Through Dynamic Mixture of ExpertsT. Y. S. S. Santosh, Tuan-Quang VuongACL 2025 · 被引用 1 次
- MTL-UE: Learning to Learn Nothing for Multi-Task LearningYi Yu, Song Xia, Siyuan Yang, Chenqi Kong 等ICML 2025
- COPER: Correlation-based Permutations for Multi-View ClusteringRan Eisenberg, Jonathan Svirsky, Ofir LindenbaumICLR 2025
- CGU-Bayes: Causal Graph Uncertainty-Guided Bayesian Inference for Domain GeneralizationNaiyu Yin, Hanjing Wang, Yue Yu, Tian Gao 等CVPR 2026
它引用的顶会 Paper27
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
相关 Paper
- Bayesian Context Aggregation for Neural ProcessesMichael Volpp, Fabian Flürenbrock, Lukas Großberger, Christian Daniel 等ICLR 2021 · 被引用 36 次
- Multi-Task Learning as a Bargaining GameAviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron 等ICML 2022 · 被引用 243 次
- Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific ParametersXiaohan Qin, Xiaoxing Wang, Junchi YanCVPR 2025
- Variational Multi-Task Learning with Gumbel-Softmax PriorsJiayi Shen, Xiantong Zhen, Marcel Worring, Ling ShaoNeurIPS 2021 · 被引用 42 次
- Task Diversity in Bayesian Federated Learning: Simultaneous Processing of Classification and RegressionJunliang Lyu, Yixuan Zhang, Xiaoling Lu, Feng ZhouKDD 2025 · 被引用 3 次
