Foundation Model Driven Appearance Extraction for Robust Multiple Object Tracking
Teng Fu, Haiyang Yu, Ke Niu, Bin Li, Xiangyang Xue
摘要
Multiple Object Tracking (MOT) is a fundamental task in computer vision. Existing methods utilize motion information or appearance information to perform object tracking. However, these algorithms still struggle with special circumstances, such as occlusion and blurring in complex scenes. Inspired by the fact that people can pinpoint objects through verbal descriptions, we explore performing long-term robust tracking using semantic features of objects. Motivated by the success of the multimodal foundation model in text-image alignment, we reconsider the appearance feature extraction module in MOT and propose a Foundation model Driven multi-object tracker (FDTracker). Specifically, we propose a two-stage trained appearance feature extractor. In the first stage, using a single image of the object as input, the model could capture the attributes of objects with the assistance of natural language instructions. In the second stage, using a sequence of images of objects as input, the model learns how to use these attributes to distinguish between different objects and connect the same object at different times. Finally, for coordinating appearance and motion information, we propose a reasonable combined strategy, which better facilitates trajectory assignment and reconnection. Extensive experiments on benchmarks demonstrate the robustness of FDTracker.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-TuningKe Niu, Zhuofan Chen, Haiyang Yu, Yuwen Chen 等NeurIPS 2025 · 被引用 8 次
- ChatReID: Open-Ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language ModelsKe Niu, Haiyang Yu, Mengyang Zhao, Teng Fu 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
相关 Paper
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan 等ICCV 2025 · 被引用 8 次
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang 等NeurIPS 2024 · 被引用 21 次
- HFF-Tracker: A Hierarchical Fine-grained Fusion Tracker for Referring Multi-Object TrackingZeyong Zhao, Yanchao Hao, Minghao Zhang, Qingbin Liu 等AAAI 2025 · 被引用 3 次
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang 等ICCV 2025 · 被引用 3 次
- Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationWangbo Zhao, Kai Wang, Xiangxiang Chu, Fuzhao Xue 等CVPR 2022 · 被引用 23 次
