MUTrack: A Memory-Aware Unified Representation Framework for Visual Tracking
Weijing Wu, Qihua Liang, Bineng Zhong, Xiaohu Tang, Yufei Tan, Ning Li, Yuanliang Xue
Abstract
Building a unified target representation that simultaneously achieves short-term adaptability and long-term stability is crucial for robust visual tracking. However, existing trackers typically face an inherent trade-off. Methods primarily relying on short-term appearance and motion cues achieve rapid adaptation, but they often struggle with long-term identity consistency. Conversely, trackers that emphasize extensive temporal context provide strong robustness, yet this approach can compromise their short-term adaptability. To bridge this gap, we propose a novel tracker, MUTrack, which comprehensively integrates both long-term and short-term memories into a unified target representation for more robust tracking. Specifically, we design a unified memory bank that stores and manages long-term memory for maintaining long-term identity consistency, and short-term memory for adapting to instantaneous appearance changes. To fully leverage the complementary nature of both long-term and short-term temporal information, we introduce a perception interaction module that dynamically fuses these memory types through deep and bidirectional interactions, enabling mutual refinement where one guides the other. This ultimately generates a highly adaptive target representation, which effectively balances adaptability to instantaneous changes with robustness against long-term identity drift. Extensive experiments on GOT10k, TrackingNet, LaSOT, LaSOT_ext, NfS, and OTB100 consistently demonstrate that MUTrack achieves SOTA performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0d4725c-fa8d-4e66-b213-104730ab3eb1Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Learning the Model Update for Siamese TrackersLichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer, Martin Danelljan et al.ICCV 2019 · 371 citations
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.AAAI 2024 · 247 citations
Related papers
- Learning Global Structure Consistency for Robust Object TrackingBi Li, Chengquan Zhang, Zhibin Hong, Xu Tang et al.ACM MM 2020 · 4 citations
- Dual-Path Temporal Decoder for End-to-End Multi-Object TrackingHyunseop Kim, Juheon Jeong, Hanul Kim, Yeong Jun KohNeurIPS 2025 · 4 citations
- Target-Aware Tracking with Long-Term Context AttentionKaijie He, Canlong Zhang, Sheng Xie, Zhixin Li et al.AAAI 2023 · 102 citations
- End-to-End Multiple Object Tracking with Dynamic Scene PerceptionRuonan Wei, Yuntao Wang, Siyan Fang, Yuehuan WangACM MM 2025 · 1 citation
- DreamTrack: Dreaming the Future for Multimodal Visual Object TrackingMingzhe Guo, Weiping Tan, Wenyu Ran, Liping Jing et al.CVPR 2025
