Argus: A Compact and Versatile Foundation Model for Vision
Weiming Zhuang, Chen Chen, Zhizhong Li, Sina Sajadmanesh, Jingtao Li, Jiabo Huang, Vikash Sehwag, Vivek Sharma, Hirotaka Shinozaki, Felan Carlo Garcia, Yihao Zhan, Naohiro Adachi
摘要
While existing vision and multi-modal foundation models can handle multiple computer vision tasks, they often suffer from significant limitations, including huge demand for data and computational resources during training and inconsistent performance across vision tasks at deployment time. To address these challenges, we introduce Argus 1 , a compact and versatile vision foundation model designed to support a wide range of vision tasks through a unified multitask architecture. Argus employs a two-stage training strategy: (i) multitask pretraining over core vision tasks with a shared backbone that includes a lightweight adapter to inject task-specific inductive biases, and (ii) scalable and efficient adaptation to new tasks by fine-tuning only the task-specific decoders. Extensive evaluations demonstrate that Argus, despite its relatively compact and trainingefficient design of merely 100M backbone parameters (only 13.6% of which are trained using 1.6M images), competes with and even surpasses much larger models. Compared to state-of-the-art foundation models, Argus not only covers a broader set of vision tasks but also matches or outperforms the models with similar sizes on 12 tasks. We expect that Argus will accelerate the real-world adoption of vision foundation models in resource-constrained scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel ManifoldZhizhong Li, Sina Sajadmanesh, Jingtao Li, Lingjuan LyuNeurIPS 2025 · 被引用 16 次
- VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution GenerationsMaitreya Patel, Jingtao Li, Weiming Zhuang, Yezhou Yang 等CVPR 2026 · 被引用 2 次
- UniCompress: Token Compression for Unified Vision-Language Understanding and GenerationZiyao Wang, Chen Chen, Jingtao Li, Weiming Zhuang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
相关 Paper
- ViM: Vision Middleware for Unified Downstream TransferringYutong Feng, Biao Gong, Jianwen Jiang, Yiliang Lv 等ICCV 2023 · 被引用 2 次
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin 等AAAI 2024 · 被引用 94 次
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等CVPR 2022 · 被引用 483 次
- 1% VS 100%: Parameter-Efficient Low Rank Adapter for Dense PredictionsDongshuo Yin, Yiran Yang, Zhechao Wang, Hongfeng Yu 等CVPR 2023
- Polyhistor: Parameter-Efficient Multi-Task Adaptation for Dense Vision TasksYen-Cheng Liu, Chih-Yao Ma, Junjiao Tian, Zijian He 等NeurIPS 2022 · 被引用 79 次
