Multiple Planar Object Tracking
Zhicheng Zhang, Shengzhe Liu, Jufeng Yang
Abstract
Tracking both location and pose of multiple planar objects (MPOT) is of great significance to numerous real-world applications. The greater degree-of-freedom of planar objects compared with common objects makes MPOT far more challenging than well-studied object tracking, especially when occlusion occurs. To address this challenging task, we are inspired by amodal perception that humans jointly track visible and invisible parts of the target, and propose a tracking framework that unifies appearance perception and occlusion reasoning. Specifically, we present a dual-branch network to track the visible part of planar objects, including vertexes and mask. Then, we develop an occlusion area localization strategy to infer the invisible part, i.e., the occluded region, followed by a two-stream attention network finally refining the prediction. To alleviate the lack of data in this field, we build the first large-scale benchmark dataset, namely MPOT-3K. It consists of 3,717 planar objects from 356 videos and contains 148,896 frames together with 687,417 annotations. The collected planar objects have 9 motion patterns and the videos are shot in 6 types of indoor and outdoor scenes. Extensive experiments demonstrate the superiority of our proposed method on the newly developed MPOT-3K as well as other two popular single planar object tracking datasets. The code and MPOT-3K dataset are released on https://zzcheng.top/MPOT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi et al.CVPR 2024 · 137 citations
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel et al.CVPR 2024 · 24 citations
- MCNet: Rethinking the Core Ingredients for Accurate and Efficient Homography EstimationHaokai Zhu, Si-Yuan Cao, Jianxin Hu, Sitong Zuo et al.CVPR 2024 · 18 citations
- LAKE-RED: Camouflaged Images Generation by Latent Background Knowledge Retrieval-Augmented DiffusionPancheng Zhao, Peng Xu, Pengda Qin, Deng-Ping Fan et al.CVPR 2024 · 14 citations
- MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution DistillationZhicheng Zhang, Pancheng Zhao, Eunil Park, Jufeng YangCVPR 2024 · 11 citations
Builds on17
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang et al.CVPR 2022 · 494 citations
- Video Instance Segmentation with a Propose-Reduce ParadigmHuaijia Lin, Ruizheng Wu, Shu Liu, Jiangbo Lu et al.ICCV 2021 · 110 citations
- Spatial Pruned Sparse Convolution for Efficient 3D Object DetectionJianhui Liu, Yukang Chen, Xiaoqing Ye, Zhuotao Tian et al.NeurIPS 2022 · 57 citations
- Planar Surface Reconstruction from Sparse ViewsLinyi Jin, Shengyi Qian, Andrew Owens, David F. FouheyICCV 2021 · 51 citations
Related papers
- Human De-Occlusion: Invisible Perception and Recovery for HumansQiang Zhou, Shiyin Wang, Yitong Wang, Zilong Huang et al.CVPR 2021
- PlanarTrack: A Large-scale Challenging Benchmark for Planar Object TrackingXinran Liu, Xiaoqiong Liu, Ziruo Yi, Xin Zhou et al.ICCV 2023 · 2 citations
- Amodal Panoptic SegmentationRohit Mohan, Abhinav ValadaCVPR 2022 · 49 citations
- PoseTrack21: A Dataset for Person Search, Multi-Object Tracking and Multi-Person Pose TrackingAndreas Doering, Di Chen, Shanshan Zhang, Bernt Schiele et al.CVPR 2022 · 47 citations
- Learning to Track with Object PermanencePavel Tokmakov, Jie Li, Wolfram Burgard, Adrien GaidonICCV 2021 · 241 citations
