ARKitTrack: A New Diverse Dataset for Tracking Using Mobile RGB-D Data
Haojie Zhao, Junsong Chen, Lijun Wang, Huchuan Lu
2023Year
5Top-tier citations
Abstract
Figure 1 . Samples from ARKitTrack. We capture both indoor and outdoor sequences (1st row) in many scenes, including zoo, market, office, square, corridor, etc. Lots of scenarios are presented in our dataset, e.g., low or high light conditions (2nd row), surrounding clutter (3rd row), out-of-plane rotation, motion blur, deformation, etc. (4th row). Besides, we annotate each frame with object masks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c111afa2-01c4-48ff-9e2a-8da5002beda7Cited by top-tier papers5
- 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular VideosZhen Xu, Zhengqin Li, Zhao Dong, Xiaowei Zhou et al.NeurIPS 2025 · 51 citations
- Don't Look Now: Audio/Haptic Guidance for 3D Scanning of LandmarksJessica Van Brummelen, Liv Piper Urwin, Oliver James Johnston, Mohamed Sayed et al.CHI 2024 · 7 citations
- X-Prompt: Multi-modal Visual Prompt for Video Object SegmentationPinxue Guo, Wanyun Li, Hao Huang, Lingyi Hong et al.ACM MM 2024 · 7 citations
- Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and BenchmarkingZhangyong Tang, Tianyang Xu, Xuefeng Zhu, Chunyang Cheng et al.ACM MM 2025 · 2 citations
- Tracking and Segmenting Anything in Any ModalityTianlu Zhang, Qiang Zhang, Guiguang Ding, Jungong HanAAAI 2026
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
Related papers
- Tracking Multiple Deformable Objects in Egocentric VideosMingzhen Huang, Xiaoxing Li, Jun Hu, Honghong Peng et al.CVPR 2023
- Learning Multi-Object Tracking and Segmentation From Automatic AnnotationsLorenzo Porzi, Markus Hofinger, Idoia Ruiz, Joan Serrat et al.CVPR 2020
- RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D VideosHongchi Xia, Yang Fu, Sifei Liu, Xiaolong WangCVPR 2024 · 14 citations
- Cubify Anything: Scaling Indoor 3D Object DetectionJustin Lazarow, David Griffiths, Gefen Kohavi, Francisco Crespo et al.CVPR 2025
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
