Vision Transformer Adapters for Generalizable Multitask Learning
Deblina Bhattacharjee, Sabine Süsstrunk, Mathieu Salzmann
摘要
We introduce the first multitasking vision transformer adapters that learn generalizable task affinities which can be applied to novel tasks and domains. Integrated into an off-the-shelf vision transformer backbone, our adapters can simultaneously solve multiple dense vision tasks in a parameter-efficient manner, unlike existing multitasking transformers that are parametrically expensive. In contrast to concurrent methods, we do not require retraining or fine-tuning whenever a new task or domain is added. We introduce a task-adapted attention mechanism within our adapter framework that combines gradient-based task similarities with attention-based ones. The learned task affinities generalize to the following settings: zero-shot task transfer, unsupervised domain adaptation, and generalization without fine-tuning to novel domains. We demonstrate that our approach outperforms not only the existing convolutional neural network-based multitasking methods but also the vision transformer-based ones. Our project page is at https://ivrl.github.io/VTAGML .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Multi-task Learning with 3D-Aware RegularizationWei-Hong Li, Steven McDonagh, Ales Leonardis, Hakan BilenICLR 2024 · 被引用 10 次
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene UnderstandingXiaoye Wang, Chen Tang, Xiangyu Yue, Wei-Hong LiCVPR 2026 · 被引用 2 次
- Argus: A Compact and Versatile Foundation Model for VisionWeiming Zhuang, Chen Chen, Zhizhong Li, Sina Sajadmanesh 等CVPR 2025
- Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained ExpertsYangyang Xu, Xi Ye, Duo SuACM MM 2025
- FAAR: Efficient Frequency-Aware Multi-Task Fine-Tuning via Automatic Rank SelectionMaxime Fontana, Michael W. Spratling, Miaojing ShiCVPR 2026
它引用的顶会 Paper30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
相关 Paper
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin 等AAAI 2024 · 被引用 94 次
- TDSS: Task Dynamic-Synergistic Skill Adaptation for Boosting Efficient and Scalable Multi-Task Learning in Dense Visual PredictionHaiming Yao, Qiyu Chen, Wei Luo, Zheng Zhang 等AAAI 2026
- Parameter-efficient Multi-task Fine-tuning for Transformers via Shared HypernetworksRabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, James HendersonACL 2021
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He 等ICLR 2023 · 被引用 204 次
- TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task LearningSeungmin Baek, Soyul Lee, Hayeon Jo, Hyesong Choi 等CVPR 2025
