DreamTrack: Dreaming the Future for Multimodal Visual Object Tracking
Mingzhe Guo, Weiping Tan, Wenyu Ran, Liping Jing, Zhipeng Zhang
Abstract
Aiming to achieve class-agnostic perception in visual object tracking, current trackers commonly formulate tracking as a one-shot detection problem with the template-matching architecture. Despite the success, severe environmental variations in long-term tracking raise challenges to generalizing the tracker in novel situations. Temporal trackers try to fix it by preserving the time-validity of target information with historical predictions, e.g., updating the template. However, solely transmitting the previous observations instead of learning from them leads to an inferior capability of understanding the tracking scenario from past experience, which is critical for the generalization in new frames. To address this issue, we reformulate temporal learning in visual tracking as a History-to-Future process and propose a novel tracking framework DreamTrack. Our Dream-Track learns the temporal dynamics from past observations to dream the future variations of the environment, which boosts the generalization with the extended future information from history. Considering the uncertainty of future variation, multimodal prediction is designed to infer the target trajectory of each possible future situation. The experiments demonstrate that our DreamTrack achieves leading performance with real-time inference speed. In particular, DreamTrack obtains SUC scores of 76.6%/87.9% on La-SOT/TrackingNet, surpassing all recent SOTA trackers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5c584ff-cc63-48c9-a034-4e654b4878c0Cited by top-tier papers11
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao et al.AAAI 2026 · 4 citations
- An Efficient Token Compression Framework for Visual Object TrackingWeijing Wu, Qihua Liang, Bineng Zhong, Haiying Xia et al.CVPR 2026 · 1 citation
- Drift-Resilient Temporal Priors for Visual TrackingYuqing Huang, Liting Lin, Weijun Zhuang, Zhenyu He et al.CVPR 2026 · 1 citation
- GOT-Edit: Geometry-Aware Generic Object Tracking via Online Model EditingShih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu LinICLR 2026 · 1 citation
- Boosting Self-Supervised Tracking with Contextual Prompts and Noise LearningYaozong Zheng, Qihua Liang, Bineng Zhong, Shuimu Zeng et al.CVPR 2026 · 1 citation
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
Related papers
- TGTrack: Temporal Generative Learning for Unified Single Object TrackingWanting Geng, Xin Chen, Chuanyu Sun, Jie Zhao et al.CVPR 2026
- Towards Generalizable Multi-Object TrackingZheng Qin, Le Wang, Sanping Zhou, Panpan Fu et al.CVPR 2024 · 21 citations
- End-to-End Multiple Object Tracking with Dynamic Scene PerceptionRuonan Wei, Yuntao Wang, Siyan Fang, Yuehuan WangACM MM 2025 · 1 citation
- Temporal Adaptive RGBT Tracking with Modality PromptHongyu Wang, Xiaotao Liu, Yifan Li, Meng Sun et al.AAAI 2024 · 92 citations
- Autoregressive Sequential Pretraining for Visual TrackingShiyi Liang, Yifan Bai, Yihong Gong, Xing WeiCVPR 2025
