iKUN: Speak to Trackers Without Retraining
Yunhao Du, Cheng Lei, Zhicheng Zhao, Fei Su
Abstract
Referring multi-object tracking (RMOT) aims to track multiple objects based on input textual descriptions. Previous works realize it by simply integrating an extra textual module into the multi-object tracker. However, they typically need to retrain the entire framework and have difficulties in optimization. In this work, we propose an insertable Knowledge Unification Network, termed iKUN, to enable communication with off-the-shelf trackers in a plug-andplay manner. Concretely, a knowledge unification module (KUM) is designed to adaptively extract visual features based on textual guidance. Meanwhile, to improve the localization accuracy, we present a neural version of Kalman filter (NKF) to dynamically adjust process noise and observation noise based on the current motion status. Moreover, to address the problem of open-set long-tail distribution of textual descriptions, a test-time similarity calibration method is proposed to refine the confidence score with pseudo frequency. Extensive experiments on Refer-KITTI dataset verify the effectiveness of our framework. Finally, to speed up the development of RMOT, we also contribute a more challenging dataset, Refer-Dance, by extending public DanceTrack dataset with motion and dressing descriptions. The codes and dataset are available at https://github.com/dyhBUPT/iKUN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72e67089-06ab-4cdd-b689-eb620bed0658Cited by top-tier papers8
- RefAV: Towards Planning-Centric Scenario MiningCainan Davidson, Deva Ramanan, Neehar PeriCVPR 2026 · 17 citations
- Cross-View Referring Multi-Object TrackingSijia Chen, En Yu, Wenbing TaoAAAI 2025 · 15 citations
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan et al.ICCV 2025 · 8 citations
- HFF-Tracker: A Hierarchical Fine-grained Fusion Tracker for Referring Multi-Object TrackingZeyong Zhao, Yanchao Hao, Minghao Zhang, Qingbin Liu et al.AAAI 2025 · 3 citations
- OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and UnderstandingTeng Fu, Mengyang Zhao, Ke Niu, Kaixin Peng et al.AAAI 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse MotionPeize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan et al.CVPR 2022 · 305 citations
Related papers
- Referring Multi-Object TrackingDongming Wu, Wencheng Han, Tiancai Wang, Xingping Dong et al.CVPR 2023
- Efficient Motion Prompt Learning for Robust Visual TrackingJie Zhao, Xin Chen, Yongsheng Yuan, Michael Felsberg et al.ICML 2025
- Foundation Model Driven Appearance Extraction for Robust Multiple Object TrackingTeng Fu, Haiyang Yu, Ke Niu, Bin Li et al.AAAI 2025 · 6 citations
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng et al.CVPR 2026
- Observation-Centric SORT: Rethinking SORT for Robust Multi-Object TrackingJinkun Cao, Jiangmiao Pang, Xinshuo Weng, Rawal Khirodkar et al.CVPR 2023
