Type-to-Track: Retrieve Any Object via Prompt-based Tracking
Pha A. Nguyen, Kha Gia Quach, Kris Kitani, Khoa Luu
摘要
One of the recent trends in vision problems is to use natural language captions to describe the objects of interest. This approach can overcome some limitations of traditional methods that rely on bounding boxes or category annotations. This paper introduces a novel paradigm for Multiple Object Tracking called Type-to-Track, which allows users to track objects in videos by typing natural language descriptions. We present a new dataset for that Grounded Multiple Object Tracking task, called GroOT, that contains videos with various types of objects and their corresponding textual captions describing their appearance and action in detail. Additionally, we introduce two new evaluation protocols and formulate evaluation metrics specifically for this task. We develop a new efficient method that models a transformer-based eMbed-ENcoDE-extRact framework (MENDER) using the third-order tensor decomposition. The experiments in five scenarios show that our MENDER approach outperforms another two-stage design in terms of accuracy and efficiency, up to 14.7% accuracy and 4 speed faster.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Language Prompt for Autonomous DrivingDongming Wu, Wencheng Han, Yingfei Liu, Tiancai Wang 等AAAI 2025 · 被引用 150 次
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren 等NeurIPS 2024 · 被引用 13 次
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 被引用 5 次
- DINTR: Tracking via Diffusion-based InterpolationPha A. Nguyen, Ngan Le, Jackson David Cothren, Alper Yilmaz 等NeurIPS 2024 · 被引用 5 次
- OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with TransformerJinyang Li, En Yu, Sijia Chen, Wenbing TaoICLR 2025
它引用的顶会 Paper19
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
相关 Paper
- Referring Multi-Object TrackingDongming Wu, Wencheng Han, Tiancai Wang, Xingping Dong 等CVPR 2023
- OmniViD: A Generative Framework for Universal Video UnderstandingJunke Wang, Dongdong Chen, Chong Luo, Bo He 等CVPR 2024 · 被引用 18 次
- Foundation Model Driven Appearance Extraction for Robust Multiple Object TrackingTeng Fu, Haiyang Yu, Ke Niu, Bin Li 等AAAI 2025 · 被引用 6 次
- Dense Video Object Captioning from Disjoint SupervisionXingyi Zhou, Anurag Arnab, Chen Sun, Cordelia SchmidICLR 2025
- End-to-End Referring Video Object Segmentation with Multimodal TransformersAdam Botach, Evgenii Zheltonozhskii, Chaim BaskinCVPR 2022 · 被引用 150 次
