Fast Template Matching and Update for Video Object Tracking and Segmentation
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Bingfeng Zhang, Yao Zhao
Abstract
In this paper, the main task we aim to tackle is the multiinstance semi-supervised video object segmentation across a sequence of frames where only the first-frame box-level ground-truth is provided. Detection-based algorithms are widely adopted to handle this task, and the challenges lie in the selection of the matching method to predict the result as well as to decide whether to update the target template using the newly predicted result. The existing methods, however, make these selections in a rough and inflexible way, compromising their performance. To overcome this limitation, we propose a novel approach which utilizes reinforcement learning to make these two decisions at the same time. Specifically, the reinforcement learning agent learns to decide whether to update the target template according to the quality of the predicted result. The choice of the matching method will be determined at the same time, based on the action history of the reinforcement learning agent. Experiments show that our method is almost 10 times faster than the previous state-of-the-art method with even higher accuracy (region similarity of 69.1% on DAVIS 2017 dataset).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 250c0d78-5be9-424e-abaa-475f9ceb5f0eCited by top-tier papers10
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 267 citations
- Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence LearningLiulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang et al.CVPR 2022 · 41 citations
- Video Object Segmentation with Dynamic Memory Networks and Adaptive Object AlignmentShuxian Liang, Xu Shen, Jianqiang Huang, Xian-Sheng HuaICCV 2021 · 28 citations
- Point-VOS: Pointing Up Video Object SegmentationSabarinath Mahadevan, Idil Esen Zulfikar, Paul Voigtlaender, Bastian LeibeCVPR 2024 · 3 citations
- RELO: Reinforcement Learning to Localize for Visual Object TrackingXin Chen, Chuanyu Sun, Jiao Xu, Houwen Peng et al.ICML 2026 · 1 citation
Builds on2
Related papers
- Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object SegmentationTianfei Zhou, Jianwu Li, Xueyi Li, Ling ShaoCVPR 2021
- A Transductive Approach for Video Object SegmentationYizhuo Zhang, Zhirong Wu, Houwen Peng, Stephen LinCVPR 2020
- Learning To Recommend Frame for Interactive Video Object Segmentation in the WildZhaoyuan Yin, Jia Zheng, Weixin Luo, Shenhan Qian et al.CVPR 2021
- Joint Inductive and Transductive Learning for Video Object SegmentationYunyao Mao, Ning Wang, Wengang Zhou, Houqiang LiICCV 2021 · 111 citations
- Learning Dynamic Network Using a Reuse Gate Function in Semi-Supervised Video Object SegmentationHyojin Park, Jayeon Yoo, Seohyeong Jeong, Ganesh Venkatesh et al.CVPR 2021
