Vision Transformer Adapters for Generalizable Multitask Learning
Deblina Bhattacharjee, Sabine Süsstrunk, Mathieu Salzmann
Abstract
We introduce the first multitasking vision transformer adapters that learn generalizable task affinities which can be applied to novel tasks and domains. Integrated into an off-the-shelf vision transformer backbone, our adapters can simultaneously solve multiple dense vision tasks in a parameter-efficient manner, unlike existing multitasking transformers that are parametrically expensive. In contrast to concurrent methods, we do not require retraining or fine-tuning whenever a new task or domain is added. We introduce a task-adapted attention mechanism within our adapter framework that combines gradient-based task similarities with attention-based ones. The learned task affinities generalize to the following settings: zero-shot task transfer, unsupervised domain adaptation, and generalization without fine-tuning to novel domains. We demonstrate that our approach outperforms not only the existing convolutional neural network-based multitasking methods but also the vision transformer-based ones. Our project page is at https://ivrl.github.io/VTAGML .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c727087-6a05-47dc-be0f-19faae5f48e8Cited by top-tier papers8
- Multi-task Learning with 3D-Aware RegularizationWei-Hong Li, Steven McDonagh, Ales Leonardis, Hakan BilenICLR 2024 · 10 citations
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene UnderstandingXiaoye Wang, Chen Tang, Xiangyu Yue, Wei-Hong LiCVPR 2026 · 2 citations
- Argus: A Compact and Versatile Foundation Model for VisionWeiming Zhuang, Chen Chen, Zhizhong Li, Sina Sajadmanesh et al.CVPR 2025
- Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained ExpertsYangyang Xu, Xi Ye, Duo SuACM MM 2025
- FAAR: Efficient Frequency-Aware Multi-Task Fine-Tuning via Automatic Rank SelectionMaxime Fontana, Michael W. Spratling, Miaojing ShiCVPR 2026
Builds on30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
Related papers
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin et al.AAAI 2024 · 94 citations
- TDSS: Task Dynamic-Synergistic Skill Adaptation for Boosting Efficient and Scalable Multi-Task Learning in Dense Visual PredictionHaiming Yao, Qiyu Chen, Wei Luo, Zheng Zhang et al.AAAI 2026
- Parameter-efficient Multi-task Fine-tuning for Transformers via Shared HypernetworksRabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, James HendersonACL 2021
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He et al.ICLR 2023 · 204 citations
- TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task LearningSeungmin Baek, Soyul Lee, Hayeon Jo, Hyesong Choi et al.CVPR 2025
