Provable Pathways: Learning Multiple Tasks over Multiple Paths
Yingcong Li, Samet Oymak
摘要
Constructing useful representations across a large number of tasks is a key requirement for sample-efficient intelligent systems. A traditional idea in multitask learning (MTL) is building a shared representation across tasks which can then be adapted to new tasks by tuning last layers. A desirable refinement of using a shared one-fits-all representation is to construct task-specific representations. To this end, recent Path-Net/muNet architectures represent individual tasks as pathways within a larger supernet. The subnetworks induced by pathways can be viewed as task-specific representations that are composition of modules within supernet's computation graph. This work explores the pathways proposal from the lens of statistical learning: We first develop novel generalization bounds for empirical risk minimization problems learning multiple tasks over multiple paths (Multipath MTL). In conjunction, we formalize the benefits of resulting multipath representation when adapting to new downstream tasks. Our bounds are expressed in terms of Gaussian complexity, lead to tangible guarantees for the class of linear representations, and provide novel insights into the quality and benefits of a multipath representation. When computation graph is a tree, Multipath MTL hierarchically clusters the tasks and builds cluster-specific representations. We provide further discussion and experiments for hierarchical MTL and rigorously identify the conditions under which Multipath MTL is provably superior to traditional MTL approaches with shallow supernets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Exploiting Shared Representations for Personalized Federated LearningLiam Collins, Hamed Hassani, Aryan Mokhtari, Sanjay ShakkottaiICML 2021 · 被引用 1,081 次
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
相关 Paper
- Deep Elastic Networks With Model Selection for Multi-Task LearningChanho Ahn, Eunwoo Kim, Songhwai OhICCV 2019 · 被引用 56 次
- Efficient Supernet Training Using Path ParallelismYing Xu, Long Cheng, Xuyi Cai, Xiaohan Ma 等HPCA 2023 · 被引用 1 次
- Efficient and Effective Multi-task Grouping via Meta Learning on Task CombinationsXiaozhuang Song, Shun Zheng, Wei Cao, James J. Q. Yu 等NeurIPS 2022 · 被引用 50 次
- MDL-NAS: A Joint Multi-domain Learning Framework for Vision TransformerShiguang Wang, Tao Xie, Jian Cheng, Xingcheng Zhang 等CVPR 2023
- Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task LearningYingru Liu, Xuewen Yang, Dongliang Xie, Xin Wang 等AAAI 2020 · 被引用 10 次
