Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
Yuxiang Lu, Shengcao Cao, Yu-Xiong Wang
Abstract
ABSTRACT Vision Foundation Models (VFMs) have demonstrated outstanding performance on numerous downstream tasks. However, due to their inherent representation biases originating from different training paradigms, VFMs exhibit advantages and disadvantages across distinct vision tasks. Although amalgamating the strengths of multiple VFMs for downstream tasks is an intuitive strategy, effectively exploiting these biases remains a significant challenge. In this paper, we propose a novel and versatile "Swiss Army Knife" (SAK) solution, which adaptively distills knowledge from a committee of VFMs to enhance multi-task learning. Unlike existing methods that use a single backbone for knowledge transfer, our approach preserves the unique representation bias of each teacher by collaborating the lightweight Teacher-Specific Adapter Path modules with the Teacher-Agnostic Stem. Through dynamic selection and combination of representations with Mixture-of-Representations Routers, our SAK is capable of synergizing the complementary strengths of multiple VFMs. Extensive experiments show that our SAK remarkably outperforms prior state of the arts in multi-task learning by 10% on the NYUD-v2 benchmark, while also providing a flexible and robust framework that can readily accommodate more advanced model designs. Project page: https://innovator-zero.github.io/SAK/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28f86299-86cd-4778-bf1b-564f23cd76a2Cited by top-tier papers8
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation ModelsSofian Chaybouti, Sanath Narayan, Yasser Dahou, Phúc H. Lê Khắc et al.CVPR 2026 · 2 citations
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene UnderstandingXiaoye Wang, Chen Tang, Xiangyu Yue, Wei-Hong LiCVPR 2026 · 2 citations
- Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense PredictionsJingdong Zhang, Hanrong Ye, Xin Li, Wenping Wang et al.ACM MM 2025 · 1 citation
- DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D TeachersMert Bülent Sariyildiz, Philippe Weinzaepfel, Thomas Lucas, Pau de Jorge et al.CVPR 2025
- PRISM: Synergizing Vision Foundation Models via Self-organized Expert SpecializationYing Tang, Dong Li, Youjia Zhang, Zikai Song et al.ICML 2026
Builds on67
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Generalizable Knowledge Distillation from Vision Foundation Models for Semantic SegmentationChonghua Lv, Dong Zhao, Shuang Wang, Dou Quan et al.CVPR 2026 · 1 citation
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic RoutingYixiao Wang, Mingxiao Huo, Zhixuan Liang, Yushi Du et al.ICLR 2026 · 5 citations
- ViM: Vision Middleware for Unified Downstream TransferringYutong Feng, Biao Gong, Jianwen Jiang, Yiliang Lv et al.ICCV 2023 · 2 citations
- All-in-One: Transferring Vision Foundation Models into Stereo MatchingJingyi Zhou, Haoyu Zhang, Jiakang Yuan, Peng Ye et al.AAAI 2025
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation ModelsBenjamin Ramtoula, Pierre-Yves Lajoie, Paul Newman, Daniele De MartiniNeurIPS 2025 · 2 citations
