Visual Prompt Multi-Modal Tracking
Jiawen Zhu, Simiao Lai, Xin Chen, Dong Wang, Huchuan Lu
摘要
Visible-modal object tracking gives rise to a series of downstream multi-modal tracking tributaries. To inherit the powerful representations of the foundation model, a natural modus operandi for multi-modal tracking is full fine-tuning on the RGB-based parameters. Albeit effective, this manner is not optimal due to the scarcity of downstream data and poor transferability, etc. In this paper, inspired by the recent success of the prompt learning in language models, we develop Visual Prompt multi-modal Tracking (ViPT), which learns the modal-relevant prompts to adapt the frozen pre-trained foundation model to various downstream multimodal tracking tasks. ViPT finds a better way to stimulate the knowledge of the RGB-based model that is pre-trained at scale, meanwhile only introducing a few trainable parameters (less than 1% of model parameters). ViPT outperforms the full fine-tuning paradigm on multiple downstream tracking tasks including RGB+Depth, RGB+Thermal, and RGB+Event tracking. Extensive experiments show the potential of visual prompt learning for multi-modal tracking, and ViPT can achieve state-of-the-art performance while satisfying parameter efficiency. Code and models are available at https://github.com/jiawen-zhu/ViPT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper72
- Bi-directional Adapter for Multimodal TrackingBing Cao, Junliang Guo, Pengfei Zhu, Qinghua HuAAAI 2024 · 被引用 153 次
- VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt LearningZiyang Luo, Nian Liu, Wangbo Zhao, Xuguang Yang 等CVPR 2024 · 被引用 96 次
- Temporal Adaptive RGBT Tracking with Modality PromptHongyu Wang, Xiaotao Liu, Yifan Li, Meng Sun 等AAAI 2024 · 被引用 92 次
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu 等CVPR 2024 · 被引用 78 次
- Generative-Based Fusion Mechanism for Multi-Modal TrackingZhangyong Tang, Tianyang Xu, Xiaojun Wu, Xuefeng Zhu 等AAAI 2024 · 被引用 78 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New BaselinePengyu Zhang, Jie Zhao, Dong Wang, Huchuan Lu 等CVPR 2022 · 被引用 225 次
相关 Paper
- Prompting for Multi-Modal TrackingJinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis 等ACM MM 2022 · 被引用 167 次
- Prompting Multi-Modal Image Segmentation with Semantic GroupingQibin HeAAAI 2024 · 被引用 21 次
- X-Prompt: Multi-modal Visual Prompt for Video Object SegmentationPinxue Guo, Wanyun Li, Hao Huang, Lingyi Hong 等ACM MM 2024 · 被引用 7 次
- SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object TrackingXiaojun Hou, Jiazheng Xing, Yijie Qian, Yaowei Guo 等CVPR 2024
- MixPrompt: Efficient Mixed Prompting for Multimodal Semantic SegmentationZhiwei Hao, Zhongyu Xiao, Jianyuan Guo, Li Shen 等NeurIPS 2025 · 被引用 1 次
