iKUN: Speak to Trackers Without Retraining
Yunhao Du, Cheng Lei, Zhicheng Zhao, Fei Su
摘要
Referring multi-object tracking (RMOT) aims to track multiple objects based on input textual descriptions. Previous works realize it by simply integrating an extra textual module into the multi-object tracker. However, they typically need to retrain the entire framework and have difficulties in optimization. In this work, we propose an insertable Knowledge Unification Network, termed iKUN, to enable communication with off-the-shelf trackers in a plug-andplay manner. Concretely, a knowledge unification module (KUM) is designed to adaptively extract visual features based on textual guidance. Meanwhile, to improve the localization accuracy, we present a neural version of Kalman filter (NKF) to dynamically adjust process noise and observation noise based on the current motion status. Moreover, to address the problem of open-set long-tail distribution of textual descriptions, a test-time similarity calibration method is proposed to refine the confidence score with pseudo frequency. Extensive experiments on Refer-KITTI dataset verify the effectiveness of our framework. Finally, to speed up the development of RMOT, we also contribute a more challenging dataset, Refer-Dance, by extending public DanceTrack dataset with motion and dressing descriptions. The codes and dataset are available at https://github.com/dyhBUPT/iKUN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- RefAV: Towards Planning-Centric Scenario MiningCainan Davidson, Deva Ramanan, Neehar PeriCVPR 2026 · 被引用 17 次
- Cross-View Referring Multi-Object TrackingSijia Chen, En Yu, Wenbing TaoAAAI 2025 · 被引用 15 次
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan 等ICCV 2025 · 被引用 8 次
- HFF-Tracker: A Hierarchical Fine-grained Fusion Tracker for Referring Multi-Object TrackingZeyong Zhao, Yanchao Hao, Minghao Zhang, Qingbin Liu 等AAAI 2025 · 被引用 3 次
- OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and UnderstandingTeng Fu, Mengyang Zhao, Ke Niu, Kaixin Peng 等AAAI 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan 等ACM MM 2022 · 被引用 314 次
- DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse MotionPeize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan 等CVPR 2022 · 被引用 305 次
相关 Paper
- Referring Multi-Object TrackingDongming Wu, Wencheng Han, Tiancai Wang, Xingping Dong 等CVPR 2023
- Efficient Motion Prompt Learning for Robust Visual TrackingJie Zhao, Xin Chen, Yongsheng Yuan, Michael Felsberg 等ICML 2025
- Foundation Model Driven Appearance Extraction for Robust Multiple Object TrackingTeng Fu, Haiyang Yu, Ke Niu, Bin Li 等AAAI 2025 · 被引用 6 次
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng 等CVPR 2026
- Observation-Centric SORT: Rethinking SORT for Robust Multi-Object TrackingJinkun Cao, Jiangmiao Pang, Xinshuo Weng, Rawal Khirodkar 等CVPR 2023
