NetTrack: Tracking Highly Dynamic Objects with a Net
Guangze Zheng, Shijie Lin, Haobo Zuo, Changhong Fu, Jia Pan
Abstract
Open-world prompt: "Osprey (Unknown) + fish (Known)" Deformation! Deformation! Referring prompt : "A large bird with brown head, wings, and white tail is predating" Fast motion! Deformation! b a Track with a Box (traditional) Fine-grained object cues Track with a Net (ours) Coarse-grained object cues Internal relationships Failure Robust Coarse-grained Fine-grained Figure 1. a The visualization of the proposed NetTrack is similar to a Net. Object dynamicity distorts the internal relationships of the object, presenting challenges for traditional coarse-grained tracking methods that rely solely on bounding boxes. While NetTrack introduces fine-grained Nets that are robust to dynamicity. b Qualitative results of NetTrack tracking highly dynamic objects under openworld tracking and referring expression comprehension settings. Dynamicity like deformation and fast motion results in drastic changes in the coarse-grained representation, while the fine-grained Nets can contract robustly. The dashed boxes represent the object position from the previous time step. c We propose a challenging benchmark named BFT, dedicated to evaluating highly dynamic object tracking with abundant scenarios shown in the external circular and diverse species shown in the central word cloud.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park et al.CVPR 2026 · 144 citations
- From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object TrackingYuqing Shao, Yuchen Yang, Rui Yu, Weilong Li et al.CVPR 2026 · 5 citations
- BlinkBud: Detecting Hazards from Behind via Sampled Monocular 3D Detection on a Single EarbudYunzhe Li, Jiajun Yan, Yuzhou Wei, Kechen Liu et al.UbiComp 2026
- FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence CompletionDian Shao, Mingfei Shi, Like LiuAAAI 2026
- JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV SystemsYifan Wang, Jian Zhao, Zhaoxin Fan, Xin Zhang et al.CVPR 2025
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- GLATrack: Global and Local Awareness for Open-Vocabulary Multiple Object TrackingGuangyao Li, Yajun Jian, Yan Yan, Hanzi WangACM MM 2024 · 3 citations
- OVTrack: Open-Vocabulary Multiple Object TrackingSiyuan Li, Tobias Fischer, Lei Ke, Henghui Ding et al.CVPR 2023
- DexTrack: Towards Generalizable Neural Tracking Control for Dexterous Manipulation from Human ReferencesXueyi Liu, Jianibieke Adalibieke, Qianwei Han, Yuzhe Qin et al.ICLR 2025
- VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object TrackingZekun Qian, Ruize Han, Junhui Hou, Linqi Song et al.ICCV 2025 · 3 citations
- Whole-Body Coordination for Dynamic Object Grasping with Legged ManipulatorsQiwei Liang, Boyang Cai, Rongyi He, Hui Li et al.AAAI 2026 · 1 citation
