Beyond Accuracy: Tracking more like Human via Visual Search
Dailing Zhang, Shiyu Hu, Xiaokun Feng, Xuchen Li, Meiqi Wu, Jing Zhang, Kaiqi Huang
摘要
Human visual search ability enables efficient and accurate tracking of an arbitrary moving target, which is a significant research interest in cognitive neuroscience. The recently proposed Central-Peripheral Dichotomy (CPD) theory sheds light on how humans effectively process visual information and track moving targets in complex environments. However, existing visual object tracking algorithms still fall short of matching human performance in maintaining tracking over time, particularly in complex scenarios requiring robust visual search skills. These scenarios often involve S patio-T emporal D iscontinuities ( i . e ., STDChallenge ), prevalent in long-term tracking and global instance tracking. To address this issue, we conduct research from a human-like modeling perspective: (1) Inspired by the CPD, we pro-pose a new tracker named CPDTrack to achieve human-like visual search ability. The central vision of CPDTrack leverages the spatio-temporal continuity of videos to introduce priors and enhance localization precision, while the peripheral vision improves global awareness and detects object movements. (2) To further evaluate and analyze STDChallenge , we create the STDChallenge Benchmark . Besides, by incorporating human subjects, we establish a human baseline, creating a high-quality environment specifically designed to assess trackers’ visual search abilities in videos across STDChallenge . (3) Our extensive experiments demonstrate
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ContextNav: Towards Agentic Multimodal In-Context LearningHonghao Fu, Yuan Ouyang, Kai-Wei Chang, Yiwei Wang 等ICLR 2026 · 被引用 14 次
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang 等ICCV 2025 · 被引用 3 次
- CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal FeaturesXiaokun Feng, Dailing Zhang, Shiyu Hu, Xuchen Li 等ICML 2025
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 被引用 746 次
- Learning Target Candidate Association to Keep Track of What Not to TrackChristoph Mayer, Martin Danelljan, Danda Pani Paudel, Luc Van GoolICCV 2021 · 被引用 356 次
相关 Paper
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo 等AAAI 2024 · 被引用 247 次
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
- GlobalTrack: A Simple and Strong Baseline for Long-Term TrackingLianghua Huang, Xin Zhao, Kaiqi HuangAAAI 2020 · 被引用 278 次
- Tracking Through Containers and Occluders in the WildBasile Van Hoorick, Pavel Tokmakov, Simon Stent, Jie Li 等CVPR 2023
- Is Tracking Really More Challenging in First Person Egocentric Vision?Matteo Dunnhofer, Zaira Manigrasso, Christian MicheloniICCV 2025 · 被引用 1 次
