Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object Tracking
Guangyao Li, Siping Zhuang, Yajun Jian, Yan Yan, Hanzi Wang
Abstract
Referring multi-object tracking (RMOT) aims to detect and track specific objects based on natural language expressions. Previous methods typically rely on sentence-level vision-language alignment, often failing to exploit finegrained linguistic cues that are crucial for distinguishing objects with similar characteristics. Notably, these cues play distinct roles at different tracking stages and should be leveraged accordingly to provide more explicit guidance. In this work, we propose DKGTrack, a novel RMOT method that enhances language comprehension for precise object tracking by decoupling language expressions into localized descriptions and motion states. To improve the accuracy of language-guided object identification, we introduce a Static Semantic Enhancement (SSE) module, which enhances region-level vision-language alignment through hierarchical cross-modal feature interaction, providing more discriminative object representations for tracking. Furthermore, we propose a Motion Perception Alignment (MPA) module that explicitly aligns object queries with motion descriptions, enabling accurate object trajectory prediction across frames. Experimental results on multiple RMOT benchmarks demonstrate the effectiveness of our method, which achieves competitive performance in challenging tracking scenarios. The code is available at https:// github.com/acyddl/DKGTrack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08f05ef7-2a84-4688-85e7-bb2e3b861a62Cited by top-tier papers2
- SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object TrackingYangkai Chen, Qiangqiang Wu, Guangyao Li, Junlong Gao et al.AAAI 2026
- Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong AgainWeize Li, Yunhao Du, Qixiang Yin, Zhicheng Zhao et al.CVPR 2026
Builds on33
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li et al.ICCV 2023 · 399 citations
- DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse MotionPeize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan et al.CVPR 2022 · 305 citations
Related papers
- Dynamic Updates for Language Adaptation in Visual-Language TrackingXiaohai Li, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.CVPR 2025
- Cross-View Referring Multi-Object TrackingSijia Chen, En Yu, Wenbing TaoAAAI 2025 · 15 citations
- Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityGuangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang et al.AAAI 2026
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang et al.ICCV 2025 · 3 citations
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng et al.CVPR 2026
