Cross-View Referring Multi-Object Tracking
Sijia Chen, En Yu, Wenbing Tao
Abstract
Referring Multi-Object Tracking (RMOT) is an important topic in the current tracking field. Its task form is to guide the tracker to track objects that match the language description. Current research mainly focuses on referring multi-object tracking under single-view, which refers to a view sequence or multiple unrelated view sequences. However, in the single-view, some appearances of objects are easily invisible, resulting in incorrect matching of objects with the language description. In this work, we propose a new task, called Cross-view Referring Multi-Object Tracking (CRMOT). It introduces the cross-view to obtain the appearances of objects from multiple views, avoiding the problem of the invisible appearances of objects in RMOT task. CRMOT is a more challenging task of accurately tracking the objects that match the language description and maintaining the identity consistency of objects in each cross-view. To advance CRMOT task, we construct a cross-view referring multi-object tracking benchmark based on CAMPUS and DIVOTrack datasets, named CRTrack. Specifically, it provides 13 different scenes and 221 language descriptions. Furthermore, we propose an end-to-end cross-view referring multi-object tracking method, named CRTracker. Extensive experiments on the CRTrack benchmark verify the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched TrainingWeifu Fu, Jinyang Li, Bin-Bin Gao, Jialin Li et al.CVPR 2026 · 3 citations
- Disentangling Instance and Scene Contexts for 3D Semantic Scene CompletionEnyu Liu, En Yu, Sijia Chen, Wenbing TaoICCV 2025 · 2 citations
- AerialMind: Towards Referring Multi-Object Tracking in UAV ScenariosChenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun et al.AAAI 2026 · 2 citations
- Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World ScenesQi Zhang, Jixuan Chen, Zhang Kaiyi, Xinquan Yu et al.CVPR 2026
- OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with TransformerJinyang Li, En Yu, Sijia Chen, Wenbing TaoICLR 2025
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search BenchmarkShuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang et al.ACM MM 2023 · 162 citations
- UCMCTrack: Multi-Object Tracking with Uniform Camera Motion CompensationKefu Yi, Kai Luo, Xiaolei Luo, Jiangui Huang et al.AAAI 2024 · 119 citations
- Self-supervised Multi-view Multi-Human Association and TrackingYiyang Gan, Ruize Han, Liqiang Yin, Wei Feng et al.ACM MM 2021 · 44 citations
- ReST: A Reconfigurable Spatial-Temporal Graph Model for Multi-Camera Multi-Object TrackingCheng-Che Cheng, Min-Xuan Qiu, Chen-Kuo Chiang, Shang-Hong LaiICCV 2023 · 39 citations
Related papers
- Referring Multi-Object TrackingDongming Wu, Wencheng Han, Tiancai Wang, Xingping Dong et al.CVPR 2023
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan et al.ICCV 2025 · 8 citations
- Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong AgainWeize Li, Yunhao Du, Qixiang Yin, Zhicheng Zhao et al.CVPR 2026
- Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object SegmentationTianming Liang, Haichao Jiang, Yuting Yang, Chaolei Tan et al.CVPR 2026 · 8 citations
- MITracker: Multi-View Integration for Visual Object TrackingMengjie Xu, Yitao Zhu, Haotian Jiang, Jiaming Li et al.CVPR 2025
