Tracking Without Re-recognition in Humans and Machines
Drew Linsley, Girik Malik, Junkyung Kim, Lakshmi Narasimhan Govindarajan, Ennio Mingolla, Thomas Serre
Abstract
Imagine trying to track one particular fruitfly in a swarm of hundreds. Higher biological visual systems have evolved to track moving objects by relying on both appearance and motion features. We investigate if state-of-the-art deep neural networks for visual tracking are capable of the same. For this, we introduce PathTracker, a synthetic visual challenge that asks human observers and machines to track a target object in the midst of identical-looking"distractor"objects. While humans effortlessly learn PathTracker and generalize to systematic variations in task design, state-of-the-art deep networks struggle. To address this limitation, we identify and model circuit mechanisms in biological brains that are implicated in tracking objects based on motion cues. When instantiated as a recurrent network, our circuit model learns to solve PathTracker with a robust visual strategy that rivals human performance and explains a significant proportion of their decision-making on the challenge. We also show that the success of this circuit model extends to object tracking in natural videos. Adding it to a transformer-based architecture for object tracking builds tolerance to visual nuisances that affect object appearance, resulting in a new state-of-the-art performance on the large-scale TrackingNet object tracking challenge. Our work highlights the importance of building artificial vision models that can help us better understand human vision and improve computer vision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Harmonizing the object recognition strategies of deep neural networks with humansThomas Fel, Ivan F. Rodriguez Rodriguez, Drew Linsley, Thomas SerreNeurIPS 2022 · 111 citations
- TarGF: Learning Target Gradient Field to Rearrange Objects without Explicit Goal SpecificationMingdong Wu, Fangwei Zhong, Yulong Xia, Hao DongNeurIPS 2022 · 25 citations
- Understanding Visual Feature Reliance through the Lens of ComplexityThomas Fel, Louis Béthune, Andrew K. Lampinen, Thomas Serre et al.NeurIPS 2024 · 20 citations
- Computing a human-like reaction time metric from stable recurrent vision modelsLore Goetschalckx, Lakshmi Narasimhan Govindarajan, Alekh Karkada Ashok, Aarit Ahuja et al.NeurIPS 2023 · 14 citations
- Beyond Accuracy: Tracking more like Human via Visual SearchDailing Zhang, Shiyu Hu, Xiaokun Feng, Xuchen Li et al.NeurIPS 2024 · 8 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistencyRobert Geirhos, Kristof Meding, Felix A. WichmannNeurIPS 2020 · 154 citations
- Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and ReasoningWeili Nie, Zhiding Yu, Lei Mao, Ankit B. Patel et al.NeurIPS 2020 · 107 citations
Related papers
- Tracking objects that change in appearance with phase synchronySabine Muzellec, Drew Linsley, Alekh Karkada Ashok, Ennio Mingolla et al.ICLR 2025
- Learning Target Candidate Association to Keep Track of What Not to TrackChristoph Mayer, Martin Danelljan, Danda Pani Paudel, Luc Van GoolICCV 2021 · 356 citations
- Recurrent neural circuits for contour detectionDrew Linsley, Junkyung Kim, Alekh Ashok, Thomas SerreICLR 2020 · 48 citations
- Video Diffusion Models Excel at Tracking Similar-Looking Objects Without SupervisionChenshuang Zhang, Kang Zhang, Joon Son Chung, In So Kweon et al.NeurIPS 2025
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman et al.ICCV 2021 · 188 citations
