MDL-NAS: A Joint Multi-domain Learning Framework for Vision Transformer
Shiguang Wang, Tao Xie, Jian Cheng, Xingcheng Zhang, Haijun Liu
Abstract
In this work, we introduce MDL-NAS, a unified framework that integrates multiple vision tasks into a manageable supernet and optimizes these tasks collectively under diverse dataset domains. MDL-NAS is storage-efficient since multiple models with a majority of shared parameters can be deposited into a single one. Technically, MDL-NAS constructs a coarse-to-fine search space, where the coarse search space offers various optimal architectures for different tasks while the fine search space provides finegrained parameter sharing to tackle the inherent obstacles of multi-domain learning. In the fine search space, we suggest two parameter sharing policies, i.e., sequential sharing policy and mask sharing policy. Compared with previous works, such two sharing policies allow for the partial sharing and non-sharing of parameters at each layer of the network, hence attaining real fine-grained parameter sharing. Finally, we present a joint-subnet search algorithm that finds the optimal architecture and sharing parameters for each task within total resource constraints, challenging the traditional practice that downstream vision tasks are typically equipped with backbone networks designed for image classification. Experimentally, we demonstrate that MDL-NAS families fitted with non-hierarchical or hierarchical transformers deliver competitive performance for all tasks compared with state-of-the-art methods while maintaining efficient storage deployment and computation. We also demonstrate that MDL-NAS allows incremental learning and evades catastrophic forgetting when generalizing to a new task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46d3ce91-4309-4b9a-bb97-3c4a82345be7Cited by top-tier papers3
- OFVL-MS: Once for Visual Localization across Multiple Indoor ScenesTao Xie, Kun Dai, Siyi Lu, Ke Wang et al.ICCV 2023 · 16 citations
- CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive NetworkTao Xie, Ke Wang, Siyi Lu, Yukun Zhang et al.ICCV 2023 · 8 citations
- Vision Transformer Neural Architecture Search for Out-of-Distribution Generalization: Benchmark and InsightsSy-Tuyen Ho, Tuan Van Vo, Somayeh Ebrahimkhani, Ngai-Man CheungNeurIPS 2024 · 5 citations
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
Related papers
- Polyhistor: Parameter-Efficient Multi-Task Adaptation for Dense Vision TasksYen-Cheng Liu, Chih-Yao Ma, Junjiao Tian, Zijian He et al.NeurIPS 2022 · 79 citations
- Task Adaptive Parameter Sharing for Multi-Task LearningMatthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran et al.CVPR 2022 · 61 citations
- Efficient Computation Sharing for Multi-Task Visual Scene UnderstandingSara Shoouri, Mingyu Yang, Zichen Fan, Hun-Seok KimICCV 2023 · 9 citations
- SUMNAS: Supernet with Unbiased Meta-Features for Neural Architecture SearchHyeonmin Ha, Ji-Hoon Kim, Semin Park, Byung-Gon ChunICLR 2022 · 5 citations
- Incremental Multi-Domain Learning with Network Latent Tensor FactorizationAdrian Bulat, Jean Kossaifi, Georgios Tzimiropoulos, Maja PanticAAAI 2020 · 33 citations
