Generating Masks from Boxes by Mining Spatio-Temporal Consistencies in Videos
Bin Zhao, Goutam Bhat, Martin Danelljan, Luc Van Gool, Radu Timofte
Abstract
Segmenting objects in videos is a fundamental computer vision task. The current deep learning based paradigm offers a powerful, but data-hungry solution. However, current datasets are limited by the cost and human effort of annotating object masks in videos. This effectively limits the performance and generalization capabilities of existing video segmentation methods. To address this issue, we explore weaker form of bounding box annotations.We introduce a method for generating segmentation masks from per-frame bounding box annotations in videos. To this end, we propose a spatio-temporal aggregation module that effectively mines consistencies in the object and background appearance across multiple frames. We use our predicted accurate masks to train video object segmentation (VOS) networks for the tracking domain, where only manual bounding box annotations are available. The additional data provides substantially better generalization performance, leading to state-of-the-art results on standard tracking benchmarks. The code and models are available at https://github.com/visionml/pytracking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and SegmentationYuanyou Xu, Zongxin Yang, Yi YangICCV 2023 · 18 citations
- Point-VOS: Pointing Up Video Object SegmentationSabarinath Mahadevan, Idil Esen Zulfikar, Paul Voigtlaender, Bastian LeibeCVPR 2024 · 3 citations
- EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric VisionYifei Cao, Yu Liu, Guolong Wang, Zhu Liu et al.AAAI 2026
Builds on4
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Probabilistic Regression for Visual TrackingMartin Danelljan, Luc Van Gool, Radu TimofteCVPR 2020
- Learning Fast and Robust Target Models for Video Object SegmentationAndreas Robinson, Felix Järemo Lawin, Martin Danelljan, Fahad Shahbaz Khan et al.CVPR 2020
Related papers
- Fast Video Object Segmentation With Temporal Aggregation Network and Dynamic Template MatchingXuhua Huang, Jiarui Xu, Yu-Wing Tai, Chi-Keung TangCVPR 2020
- Frame-to-Frame Aggregation of Active Regions in Web Videos for Weakly Supervised Semantic SegmentationJungbeom Lee, Eunji Kim, Sungmin Lee, Jangho Lee et al.ICCV 2019 · 45 citations
- Towards Robust Video Object Segmentation with Adaptive Object CalibrationXiaohao Xu, Jinglu Wang, Xiang Ming, Yan LuACM MM 2022 · 21 citations
- SRNet: Spatial Relation Network for Efficient Single-stage Instance Segmentation in VideosXiaowen Ying, Xin Li, Mooi Choo ChuahACM MM 2021 · 4 citations
- Query-Memory Re-Aggregation for Weakly-supervised Video Object SegmentationFanchao Lin, Hongtao Xie, Yan Li, Yongdong ZhangAAAI 2021 · 26 citations
