Learning To Recommend Frame for Interactive Video Object Segmentation in the Wild
Zhaoyuan Yin, Jia Zheng, Weixin Luo, Shenhan Qian, Hanling Zhang, Shenghua Gao
Abstract
This paper proposes a framework for the interactive video object segmentation (VOS) in the wild where users can choose some frames for annotations iteratively. Then, based on the user annotations, a segmentation algorithm refines the masks. The previous interactive VOS paradigm selects the frame with some worst evaluation metric, and the ground truth is required for calculating the evaluation metric, which is impractical in the testing phase. In contrast, in this paper, we advocate that the frame with the worst evaluation metric may not be exactly the most valuable frame that leads to the most performance improvement across the video. Thus, we formulate the frame selection problem in the interactive VOS as a Markov Decision Process, where an agent is learned to recommend the frame under a deep reinforcement learning framework. The learned agent can automatically determine the most valuable frame, making the interactive setting more practical in the wild. Experimental results on the public datasets show the effectiveness of our learned agent without any changes to the underlying VOS algorithms. Our data, code, and models are available at https://github.com/svip-lab/IVOS-W .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fa05daf-b4ca-4bfc-97b5-d6b886cf9f3aCited by top-tier papers3
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 267 citations
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
- XMem++: Production-level Video Segmentation From Few Annotated FramesMaksym Bekuzarov, Ariana Bermudez, Joon-Young Lee, Hao LiICCV 2023 · 69 citations
Builds on5
- DMM-Net: Differentiable Mask-Matching Network for Video Object SegmentationXiaohui Zeng, Renjie Liao, Li Gu, Yuwen Xiong et al.ICCV 2019 · 78 citations
- StartNet: Online Detection of Action Start in Untrimmed VideosMingfei Gao, Mingze Xu, Larry Davis, Richard Socher et al.ICCV 2019 · 56 citations
- Patchwork: A Patch-Wise Attention Network for Efficient Object Detection and Segmentation in Video StreamsYuning ChaiICCV 2019 · 32 citations
- Memory Aggregation Networks for Efficient Interactive Video Object SegmentationJiaxu Miao, Yunchao Wei, Yi YangCVPR 2020
- Progressive Relation Learning for Group Activity RecognitionGuyue Hu, Bo Cui, Yuan He, Shan YuCVPR 2020
Related papers
- Fast Template Matching and Update for Video Object Tracking and SegmentationMingjie Sun, Jimin Xiao, Eng Gee Lim, Bingfeng Zhang et al.CVPR 2020
- Iteratively-Refined Interactive 3D Medical Image Segmentation With Multi-Agent Reinforcement LearningXuan Liao, Wenhao Li, Qisen Xu, Xiangfeng Wang et al.CVPR 2020
- Dynamic Face Video Segmentation via Reinforcement LearningYujiang Wang, Mingzhi Dong, Jie Shen, Yang Wu et al.CVPR 2020
- Guided Interactive Video Object Segmentation Using Reliability-Based Attention MapsYuk Heo, Yeong Jun Koh, Chang-Su KimCVPR 2021
- Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware FusionHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangCVPR 2021
