Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
Yuxiang Lu, Shengcao Cao, Yu-Xiong Wang
摘要
ABSTRACT Vision Foundation Models (VFMs) have demonstrated outstanding performance on numerous downstream tasks. However, due to their inherent representation biases originating from different training paradigms, VFMs exhibit advantages and disadvantages across distinct vision tasks. Although amalgamating the strengths of multiple VFMs for downstream tasks is an intuitive strategy, effectively exploiting these biases remains a significant challenge. In this paper, we propose a novel and versatile "Swiss Army Knife" (SAK) solution, which adaptively distills knowledge from a committee of VFMs to enhance multi-task learning. Unlike existing methods that use a single backbone for knowledge transfer, our approach preserves the unique representation bias of each teacher by collaborating the lightweight Teacher-Specific Adapter Path modules with the Teacher-Agnostic Stem. Through dynamic selection and combination of representations with Mixture-of-Representations Routers, our SAK is capable of synergizing the complementary strengths of multiple VFMs. Extensive experiments show that our SAK remarkably outperforms prior state of the arts in multi-task learning by 10% on the NYUD-v2 benchmark, while also providing a flexible and robust framework that can readily accommodate more advanced model designs. Project page: https://innovator-zero.github.io/SAK/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation ModelsSofian Chaybouti, Sanath Narayan, Yasser Dahou, Phúc H. Lê Khắc 等CVPR 2026 · 被引用 2 次
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene UnderstandingXiaoye Wang, Chen Tang, Xiangyu Yue, Wei-Hong LiCVPR 2026 · 被引用 2 次
- Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense PredictionsJingdong Zhang, Hanrong Ye, Xin Li, Wenping Wang 等ACM MM 2025 · 被引用 1 次
- DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D TeachersMert Bülent Sariyildiz, Philippe Weinzaepfel, Thomas Lucas, Pau de Jorge 等CVPR 2025
- PRISM: Synergizing Vision Foundation Models via Self-organized Expert SpecializationYing Tang, Dong Li, Youjia Zhang, Zikai Song 等ICML 2026
它引用的顶会 Paper67
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Generalizable Knowledge Distillation from Vision Foundation Models for Semantic SegmentationChonghua Lv, Dong Zhao, Shuang Wang, Dou Quan 等CVPR 2026 · 被引用 1 次
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic RoutingYixiao Wang, Mingxiao Huo, Zhixuan Liang, Yushi Du 等ICLR 2026 · 被引用 5 次
- ViM: Vision Middleware for Unified Downstream TransferringYutong Feng, Biao Gong, Jianwen Jiang, Yiliang Lv 等ICCV 2023 · 被引用 2 次
- All-in-One: Transferring Vision Foundation Models into Stereo MatchingJingyi Zhou, Haoyu Zhang, Jiakang Yuan, Peng Ye 等AAAI 2025
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation ModelsBenjamin Ramtoula, Pierre-Yves Lajoie, Paul Newman, Daniele De MartiniNeurIPS 2025 · 被引用 2 次
