Structure Matters: Revisiting Boundary Refinement in Video Object Segmentation
Guanyi Qin, Ziyue Wang, Daiyun Shen, Haofeng Liu, Hantao Zhou, Junde Wu, Runze Hu, Yueming Jin
Abstract
Given an object mask, Semi-supervised Video Object Segmentation (SVOS) technique aims to track and segment the object across video frames, serving as a fundamental task in computer vision. Although recent memory-based methods demonstrate potential, they often struggle with scenes involving occlusion, particularly in handling object interactions and high feature similarity. To address these issues and meet the real-time processing requirements of downstream applications, in this paper, we propose a novel bOundary Amendment video object Segmentation method with Inherent Structure refinement, hereby named OASIS. Specifically, a lightweight structure refinement module is proposed to enhance segmentation accuracy. With the fusion of rough edge priors captured by the Canny filter and stored object features, the module can generate an object-level structure map and refine the representations by highlighting boundary features. Evidential learning for uncertainty estimation is introduced to further address challenges in occluded regions. The proposed method, OASIS, maintains an efficient design, yet extensive experiments on challenging benchmarks demonstrate its superior performance and competitive inference speed compared to other state-of-the-art methods, i.e., achieving the F values of 91.6 (vs. 89.7 on DAVIS-17 validation set) and G values of 86.6 (vs. 86.2 on YouTubeVOS 2019 validation set) while maintaining a competitive speed of 48 FPS on DAVIS. Checkpoints, logs, and code will be available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement LearningSenqiao Yang, Junyi Li, Xin Lai, Jinming Wu et al.NeurIPS 2025 · 43 citations
- DR.Experts: Differential Refinement of Distortion-Aware Experts for Blind Image Quality AssessmentBohan Fu, Guanyi Qin, Fazhan Zhang, Zihao Huang et al.AAAI 2026 · 1 citation
Builds on30
- EGNet: Edge Guidance Network for Salient Object DetectionJiaxing Zhao, Jiang-Jiang Liu, Deng-Ping Fan, Yang Cao et al.ICCV 2019 · 1,054 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- Stacked Cross Refinement Network for Edge-Aware Salient Object DetectionZhe Wu, Li Su, Qingming HuangICCV 2019 · 374 citations
Related papers
- Towards Robust Video Object Segmentation with Adaptive Object CalibrationXiaohao Xu, Jinglu Wang, Xiang Ming, Yan LuACM MM 2022 · 21 citations
- SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-MaximizationZhihui Lin, Tianyu Yang, Maomao Li, Ziyu Wang et al.CVPR 2022 · 42 citations
- Per-Clip Video Object SegmentationKwanyong Park, Sanghyun Woo, Seoung Wug Oh, In So Kweon et al.CVPR 2022 · 45 citations
- Video Object Segmentation with Dynamic Memory Networks and Adaptive Object AlignmentShuxian Liang, Xu Shen, Jianqiang Huang, Xian-Sheng HuaICCV 2021 · 28 citations
- Learning Position and Target Consistency for Memory-Based Video Object SegmentationLi Hu, Peng Zhang, Bang Zhang, Pan Pan et al.CVPR 2021
