Learning to Learn Better for Video Object Segmentation
Meng Lan, Jing Zhang, Lefei Zhang, Dacheng Tao
Abstract
Recently, the joint learning framework (JOINT) integrates matching based transductive reasoning and online inductive learning to achieve accurate and robust semi-supervised video object segmentation (SVOS). However, using the mask embedding as the label to guide the generation of target features in the two branches may result in inadequate target representation and degrade the performance. Besides, how to reasonably fuse the target features in the two different branches rather than simply adding them together to avoid the adverse effect of one dominant branch has not been investigated. In this paper, we propose a novel framework that emphasizes Learning to Learn Better (LLB) target features for SVOS, termed LLB, where we design the discriminative label generation module (DLGM) and the adaptive fusion module to address these issues. Technically, the DLGM takes the background-filtered frame instead of the target mask as input and adopts a lightweight encoder to generate the target features, which serves as the label of the online few-shot learner and the value of the decoder in the transformer to guide the two branches to learn more discriminative target representation. The adaptive fusion module maintains a learnable gate for each branch, which reweighs the element-wise feature representation and allows an adaptive amount of target information in each branch flowing to the fused target feature, thus preventing one branch from being dominant and making the target feature more robust to distractor. Extensive experiments on public benchmarks show that our proposed LLB method achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87aa366f-8abc-42df-9a55-e7957d59b610Cited by top-tier papers5
- XMem++: Production-level Video Segmentation From Few Annotated FramesMaksym Bekuzarov, Ariana Bermudez, Joon-Young Lee, Hao LiICCV 2023 · 69 citations
- Multi-Task Learning with Knowledge Distillation for Dense PredictionYangyang Xu, Yibo Yang, Lefei ZhangICCV 2023 · 18 citations
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 4 citations
- Structure Matters: Revisiting Boundary Refinement in Video Object SegmentationGuanyi Qin, Ziyue Wang, Daiyun Shen, Haofeng Liu et al.ICCV 2025 · 1 citation
- Token Contrast for Weakly-Supervised Semantic SegmentationLixiang Ru, Heliang Zheng, Yibing Zhan, Bo DuCVPR 2023
Builds on18
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region RefinementYongqing Liang, Xin Li, Navid H. Jafari, Jim ChenNeurIPS 2020 · 192 citations
Related papers
- Joint Inductive and Transductive Learning for Video Object SegmentationYunyao Mao, Ning Wang, Wengang Zhou, Houqiang LiICCV 2021 · 111 citations
- Siamese Network with Interactive Transformer for Video Object SegmentationMeng Lan, Jing Zhang, Fengxiang He, Lefei ZhangAAAI 2022 · 41 citations
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li et al.CVPR 2023
- Learning Position and Target Consistency for Memory-Based Video Object SegmentationLi Hu, Peng Zhang, Bang Zhang, Pan Pan et al.CVPR 2021
- Beyond Pixel and Object: Part Feature as Reference for Few-Shot Video Object SegmentationNaisong Luo, Guoxin Xiong, Tianzhu ZhangAAAI 2025
