Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation
Ho Kei Cheng, Yu-Wing Tai, Chi-Keung Tang
Abstract
This paper presents a simple yet effective approach to modeling space-time correspondences in the context of video object segmentation. Unlike most existing approaches, we establish correspondences directly between frames without re-encoding the mask features for every object, leading to a highly efficient and robust framework. With the correspondences, every node in the current query frame is inferred by aggregating features from the past in an associative fashion. We cast the aggregation process as a voting problem and find that the existing inner-product affinity leads to poor use of memory with a small (fixed) subset of memory nodes dominating the votes, regardless of the query. In light of this phenomenon, we propose using the negative squared Euclidean distance instead to compute the affinities. We validated that every memory node now has a chance to contribute, and experimentally showed that such diversified voting is beneficial to both memory efficiency and inference accuracy. The synergy of correspondence networks and diversified voting works exceedingly well, achieves new state-of-the-art results on both DAVIS and YouTubeVOS datasets while running significantly faster at 20+ FPS for multiple objects without bells and whistles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d27a6e6-053d-41aa-bb13-5d3a5b27e0a8Cited by top-tier papers104
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 267 citations
- Decoupling Features in Hierarchical Propagation for Video Object SegmentationZongxin Yang, Yi YangNeurIPS 2022 · 243 citations
- Tracking Anything with Decoupled Video SegmentationHo Kei Cheng, Seoung Wug Oh, Brian L. Price, Alexander G. Schwing et al.ICCV 2023 · 240 citations
- Language as Queries for Referring Video Object SegmentationJiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan et al.CVPR 2022 · 143 citations
- LVOS: A Benchmark for Long-term Video Object SegmentationLingyi Hong, Wenchao Chen, Zhongying Liu, Wei Zhang et al.ICCV 2023 · 89 citations
Builds on23
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou et al.ICCV 2019 · 1,404 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Towards High-Resolution Salient Object DetectionYi Zeng, Pingping Zhang, Zhe Lin, Jianming Zhang et al.ICCV 2019 · 232 citations
- RANet: Ranking Attention Network for Fast Video Object SegmentationZiqin Wang, Jun Xu, Li Liu, Fan Zhu et al.ICCV 2019 · 217 citations
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 208 citations
Related papers
- Per-Clip Video Object SegmentationKwanyong Park, Sanghyun Woo, Seoung Wug Oh, In So Kweon et al.CVPR 2022 · 45 citations
- Boosting Video Object Segmentation via Space-Time Correspondence LearningYurong Zhang, Liulei Li, Wenguan Wang, Rong Xie et al.CVPR 2023
- Memory Aggregation Networks for Efficient Interactive Video Object SegmentationJiaxu Miao, Yunchao Wei, Yi YangCVPR 2020
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li et al.CVPR 2023
- SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-MaximizationZhihui Lin, Tianyu Yang, Maomao Li, Ziyu Wang et al.CVPR 2022 · 42 citations
