Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation
Xiangyu Zheng, Songcheng He, Wanyun Li, Xiaoqiang Li, Wei Zhang
Abstract
Unsupervised Video Object Segmentation (UVOS) aims to predict pixel-level masks for the most salient objects in videos without any prior annotations. While memory mechanisms have been proven critical in various video segmentation paradigms, their application in UVOS yield only marginal performance gains despite sophisticated design. Our analysis reveals a simple but fundamental flaw in existing methods: over-reliance on memorizing high-level semantic features. UVOS inherently suffers from the deficiency of lacking fine-grained information due to the absence of pixel-level prior knowledge. Consequently, memory design relying solely on high-level features, which predominantly capture abstract semantic cues, is insufficient to generate precise predictions. To resolve this fundamental issue, we propose a novel hierarchical memory architecture to incorporate both shallow- and high-level features for memory, which leverages the complementary benefits of pixel and semantic information. Furthermore, to balance the simultaneous utilization of the pixel and semantic memory features, we propose a heterogeneous interaction mechanism to perform pixel-semantic mutual interactions, which explicitly considers their inherent feature discrepancies. Through the design of Pixel-guided Local Alignment Module (PLAM) and Semantic-guided Global Integration Module (SGIM), we achieve delicate integration of the fine-grained details in shallow-level memory and the semantic representations in high-level memory. Our Hierarchical Memory with Heterogeneous Interaction Network (HMHI-Net) consistently achieves state-of-the-art performance across all UVOS and video saliency detection benchmarks. Moreover, HMHI-Net consistently exhibits high performance across different backbones, further demonstrating its superiority and robustness. Project page: https://github.com/ZhengxyFlow/HMHI-Net .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b740f7fb-e6f3-40c6-9fe0-bfcb6753acaaBuilds on27
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic SegmentationLukas Hoyer, Dengxin Dai, Luc Van GoolCVPR 2022 · 562 citations
- Semi-Supervised Semantic Segmentation Using Unreliable Pseudo-LabelsYuchao Wang, Haochen Wang, Yujun Shen, Jingjing Fei et al.CVPR 2022 · 448 citations
Related papers
- Hierarchical Memory Matching Network for Video Object SegmentationHongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee et al.ICCV 2021 · 126 citations
- Dual Prototype Attention for Unsupervised Video Object SegmentationSuhwan Cho, Minhyeok Lee, Seunghoon Lee, Dogyoon Lee et al.CVPR 2024
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu et al.ACM MM 2023 · 14 citations
- Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationJiahua Dong, Hui Yin, Wenqi Liang, Hanbin Zhao et al.ICCV 2025
- VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAEHaonan Yu, Wei XuICLR 2024 · 1 citation
