Alignment Before Aggregation: Trajectory Memory Retrieval Network for Video Object Segmentation
Rui Sun, Yuan Wang, Huayu Mai, Tianzhu Zhang, Feng Wu
Abstract
Memory-based methods in semi-supervised video object segmentation task achieve competitive performance by performing dense matching between query and memory frames. However, most of the existing methods neglect the fact that videos carry rich temporal information yet redundant spatial information. In this case, direct pixel-level global matching will lead to ambiguous correspondences. In this work, we reconcile the inherent tension of spatial and temporal information to retrieve memory frame information along the object trajectory, and propose a novel and coherent Trajectory Memory Retrieval Network (TMRN) to equip with the trajectory information, including a spatial alignment module and a temporal aggregation module. The proposed TMRN enjoys several merits. First, TMRN is empowered to characterize the temporal correspondence which is in line with the nature of video in a data-driven manner. Second, we elegantly customize the spatial alignment module by coupling SVD initialization with agent-level correlation for representative agent construction and rectifying false matches caused by direct pairwise pixel-level correlation, respectively. Extensive experimental results on challenging benchmarks including DAVIS 2017 validation / test and Youtube-VOS 2018 / 2019 demonstrate that our TMRN, as a general plugin module, achieves consistent improvements over several leading methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- RankMatch: Exploring the Better Consistency Regularization for Semi-Supervised Semantic SegmentationHuayu Mai, Rui Sun, Tianzhu Zhang, Feng WuCVPR 2024 · 48 citations
- DAW: Exploring the Better Weighting Function for Semi-supervised Semantic SegmentationRui Sun, Huayu Mai, Tianzhu Zhang, Feng WuNeurIPS 2023 · 40 citations
- Focus on Query: Adversarial Mining Transformer for Few-Shot SegmentationYuan Wang, Naisong Luo, Tianzhu ZhangNeurIPS 2023 · 29 citations
- Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic SegmentationYuan Wang, Rui Sun, Naisong Luo, Yuwen Pan et al.CVPR 2024 · 13 citations
- Pay Attention to Target: Relation-Aware Temporal Consistency for Domain Adaptive Video Semantic SegmentationHuayu Mai, Rui Sun, Yuan Wang, Tianzhu Zhang et al.AAAI 2024 · 12 citations
Builds on31
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- Decoupling Features in Hierarchical Propagation for Video Object SegmentationZongxin Yang, Yi YangNeurIPS 2022 · 243 citations
- Towards High-Resolution Salient Object DetectionYi Zeng, Pingping Zhang, Zhe Lin, Jianming Zhang et al.ICCV 2019 · 232 citations
Related papers
- Video Object Segmentation with Dynamic Memory Networks and Adaptive Object AlignmentShuxian Liang, Xu Shen, Jianqiang Huang, Xian-Sheng HuaICCV 2021 · 28 citations
- Hierarchical Memory Matching Network for Video Object SegmentationHongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee et al.ICCV 2021 · 126 citations
- Dual Temporal Memory Network for Efficient Video Object SegmentationKaihua Zhang, Long Wang, Dong Liu, Bo Liu et al.ACM MM 2020 · 16 citations
- SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-MaximizationZhihui Lin, Tianyu Yang, Maomao Li, Ziyu Wang et al.CVPR 2022 · 42 citations
- Efficient Regional Memory Network for Video Object SegmentationHaozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang et al.CVPR 2021
