Prompting for Multi-Modal Tracking
Jinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis, Jingkuan Song
摘要
Multi-modal tracking gains attention due to its ability to be more accurate and robust in complex scenarios compared to traditional RGB-based tracking. Its key lies in how to fuse multi-modal data and reduce the gap between modalities. However, multi-modal tracking still severely suffers from data deficiency, thus resulting in the insufficient learning of fusion modules. Instead of building such a fusion module, in this paper, we provide a new perspective on multi-modal tracking by attaching importance to the multi-modal visual prompts. We design a novel multi-modal prompt tracker (ProTrack), which can transfer the multi-modal inputs to a single modality by the prompt paradigm. By best employing the tracking ability of pre-trained RGB trackers learning at scale, our ProTrack can achieve high-performance multi-modal tracking by only altering the inputs, even without any extra training on multi-modal data. Extensive experiments on 5 benchmark datasets demonstrate the effectiveness of the proposed ProTrack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Bi-directional Adapter for Multimodal TrackingBing Cao, Junliang Guo, Pengfei Zhu, Qinghua HuAAAI 2024 · 被引用 153 次
- Temporal Adaptive RGBT Tracking with Modality PromptHongyu Wang, Xiaotao Liu, Yifan Li, Meng Sun 等AAAI 2024 · 被引用 92 次
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu 等CVPR 2024 · 被引用 78 次
- Generative-Based Fusion Mechanism for Multi-Modal TrackingZhangyong Tang, Tianyang Xu, Xiaojun Wu, Xuefeng Zhu 等AAAI 2024 · 被引用 78 次
- Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingXiantao Hu, Ying Tai, Xu Zhao, Chen Zhao 等AAAI 2025 · 被引用 65 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- Learning Target Candidate Association to Keep Track of What Not to TrackChristoph Mayer, Martin Danelljan, Danda Pani Paudel, Luc Van GoolICCV 2021 · 被引用 356 次
相关 Paper
- Visual Prompt Multi-Modal TrackingJiawen Zhu, Simiao Lai, Xin Chen, Dong Wang 等CVPR 2023
- OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient TuningLingyi Hong, Shilin Yan, Renrui Zhang, Wanyun Li 等CVPR 2024
- SMStracker: Tri-Path Score Mask Sigma Fusion for Multi-Modal TrackingSixian Chan, Zedong Li, Wenhao Li, Shijian Lu 等ICCV 2025 · 被引用 3 次
- SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object TrackingXiaojun Hou, Jiazheng Xing, Yijie Qian, Yaowei Guo 等CVPR 2024
- Efficient Motion Prompt Learning for Robust Visual TrackingJie Zhao, Xin Chen, Yongsheng Yuan, Michael Felsberg 等ICML 2025
