RGBD1K: A Large-Scale Dataset and Benchmark for RGB-D Object Tracking
Xuefeng Zhu, Tianyang Xu, Zhangyong Tang, Zucheng Wu, Haodong Liu, Xiao Yang, Xiao-Jun Wu, Josef Kittler
Abstract
RGB-D object tracking has attracted considerable attention recently, achieving promising performance thanks to the symbiosis between visual and depth channels. However, given a limited amount of annotated RGB-D tracking data, most state-of-the-art RGB-D trackers are simple extensions of high-performance RGB-only trackers, without fully exploiting the underlying potential of the depth channel in the offline training stage. To address the dataset deficiency issue, a new RGB-D dataset named RGBD1K is released in this paper. The RGBD1K contains 1,050 sequences with about 2.5M frames in total. To demonstrate the benefits of training on a larger RGB-D data set in general, and RGBD1K in particular, we develop a transformer-based RGB-D tracker, named SPT, as a baseline for future visual object tracking studies using the new dataset. The results, of extensive experiments using the SPT tracker demonstrate the potential of the RGBD1K dataset to improve the performance of RGB-D tracking, inspiring future developments of effective tracker designs. The dataset and codes will be available on the project homepage: https://github.com/xuefeng-zhu5/RGBD1K.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44f8aeec-362f-4c94-9464-3f69844188dcCited by top-tier papers11
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu et al.CVPR 2024 · 78 citations
- Generative-Based Fusion Mechanism for Multi-Modal TrackingZhangyong Tang, Tianyang Xu, Xiaojun Wu, Xuefeng Zhu et al.AAAI 2024 · 78 citations
- Breaking Modality Gap in RGBT Tracking: Coupled Knowledge DistillationAndong Lu, Jiacong Zhao, Chenglong Li, Yun Xiao et al.ACM MM 2024 · 15 citations
- XTrack: Multimodal Training Boosts RGB-X Video Object TrackersYuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou et al.ICCV 2025 · 10 citations
- What You Have is What You Track: Adaptive and Robust Multimodal TrackingYuedong Tan, Jiawei Shao, Eduard Zamfir, Ruanjun Li et al.ICCV 2025 · 5 citations
Builds on8
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Target Candidate Association to Keep Track of What Not to TrackChristoph Mayer, Martin Danelljan, Danda Pani Paudel, Luc Van GoolICCV 2021 · 356 citations
- Joint Group Feature Selection and Discriminative Filter Learning for Robust Visual Object TrackingTianyang Xu, Zhenhua Feng, Xiao-Jun Wu, Josef KittlerICCV 2019 · 182 citations
- DepthTrack: Unveiling the Power of RGBD TrackingSong Yan, Jinyu Yang, Jani Käpylä, Feng Zheng et al.ICCV 2021 · 114 citations
- CDTB: A Color and Depth Visual Object Tracking Dataset and BenchmarkAlan Lukezic, Ugur Kart, Jani Käpylä, Ahmed Durmush et al.ICCV 2019 · 79 citations
Related papers
- GSOT3D: Towards Generic 3D Single Object Tracking in the WildYifan Jiao, Yunhao Li, Junhua Ding, Qing Yang et al.ICCV 2025
- MUST: The First Dataset and Unified Framework for Multispectral UAV Single Object TrackingHaolin Qin, Tingfa Xu, Tianhao Li, Zhenxiang Chen et al.CVPR 2025
- Resource-Efficient RGBD Aerial TrackingJinyu Yang, Shang Gao, Zhe Li, Feng Zheng et al.CVPR 2023
- DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationBowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu et al.ICLR 2024 · 110 citations
- Cross-Modal Object Tracking: Modality-Aware Representations and a Unified BenchmarkChenglong Li, Tianhao Zhu, Lei Liu, Xiaonan Si et al.AAAI 2022 · 12 citations
