Hierarchical Memory Matching Network for Video Object Segmentation
Hongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee, Suhyeon Lee, Euntai Kim
Abstract
We present Hierarchical Memory Matching Network (HMMN) for semi-supervised video object segmentation. Based on a recent memory-based method [33], we propose two advanced memory read modules that enable us to perform memory reading in multiple scales while exploiting temporal smoothness. We first propose a kernel guided memory matching module that replaces the non-local dense memory read, commonly adopted in previous memory-based methods. The module imposes the temporal smoothness constraint in the memory read, leading to accurate memory retrieval. More importantly, we introduce a hierarchical memory matching scheme and propose a top-k guided memory matching module in which memory read on a fine-scale is guided by that on a coarse-scale. With the module, we perform memory read in multiple scales efficiently and leverage both high-level semantic and low-level fine-grained memory features to predict detailed object masks. Our network achieves state-of-the-art performance on the validation sets of DAVIS 2016/2017 (90.8% and 84.7%) and YouTube-VOS 2018/2019 (82.6% and 82.5%), and test-dev set of DAVIS 2017 (78.6%). The source code and model are available online: https://github.com/Hongje/HMMN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers36
- Decoupling Features in Hierarchical Propagation for Video Object SegmentationZongxin Yang, Yi YangNeurIPS 2022 · 243 citations
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang et al.CVPR 2024 · 95 citations
- WildNet: Learning Domain Generalized Semantic Segmentation from the WildSuhyeon Lee, Hongje Seong, Seongwon Lee, Euntai KimCVPR 2022 · 95 citations
- LVOS: A Benchmark for Long-term Video Object SegmentationLingyi Hong, Wenchao Chen, Zhongying Liu, Wei Zhang et al.ICCV 2023 · 89 citations
- Recurrent Dynamic Embedding for Video Object SegmentationMingxing Li, Li Hu, Zhiwei Xiong, Bang Zhang et al.CVPR 2022 · 80 citations
Builds on16
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- RANet: Ranking Attention Network for Fast Video Object SegmentationZiqin Wang, Jun Xu, Li Liu, Fan Zhu et al.ICCV 2019 · 217 citations
- Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region RefinementYongqing Liang, Xin Li, Navid H. Jafari, Jim ChenNeurIPS 2020 · 192 citations
- AGSS-VOS: Attention Guided Single-Shot Video Object SegmentationHuaijia Lin, Xiaojuan Qi, Jiaya JiaICCV 2019 · 94 citations
Related papers
- Video Object Segmentation with Dynamic Memory Networks and Adaptive Object AlignmentShuxian Liang, Xu Shen, Jianqiang Huang, Xian-Sheng HuaICCV 2021 · 28 citations
- Per-Clip Video Object SegmentationKwanyong Park, Sanghyun Woo, Seoung Wug Oh, In So Kweon et al.CVPR 2022 · 45 citations
- Efficient Regional Memory Network for Video Object SegmentationHaozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang et al.CVPR 2021
- Alignment Before Aggregation: Trajectory Memory Retrieval Network for Video Object SegmentationRui Sun, Yuan Wang, Huayu Mai, Tianzhu Zhang et al.ICCV 2023 · 12 citations
- Dual Temporal Memory Network for Efficient Video Object SegmentationKaihua Zhang, Long Wang, Dong Liu, Bo Liu et al.ACM MM 2020 · 16 citations
