Learning to Learn Better for Video Object Segmentation
Meng Lan, Jing Zhang, Lefei Zhang, Dacheng Tao
摘要
Recently, the joint learning framework (JOINT) integrates matching based transductive reasoning and online inductive learning to achieve accurate and robust semi-supervised video object segmentation (SVOS). However, using the mask embedding as the label to guide the generation of target features in the two branches may result in inadequate target representation and degrade the performance. Besides, how to reasonably fuse the target features in the two different branches rather than simply adding them together to avoid the adverse effect of one dominant branch has not been investigated. In this paper, we propose a novel framework that emphasizes Learning to Learn Better (LLB) target features for SVOS, termed LLB, where we design the discriminative label generation module (DLGM) and the adaptive fusion module to address these issues. Technically, the DLGM takes the background-filtered frame instead of the target mask as input and adopts a lightweight encoder to generate the target features, which serves as the label of the online few-shot learner and the value of the decoder in the transformer to guide the two branches to learn more discriminative target representation. The adaptive fusion module maintains a learnable gate for each branch, which reweighs the element-wise feature representation and allows an adaptive amount of target information in each branch flowing to the fused target feature, thus preventing one branch from being dominant and making the target feature more robust to distractor. Extensive experiments on public benchmarks show that our proposed LLB method achieves state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- XMem++: Production-level Video Segmentation From Few Annotated FramesMaksym Bekuzarov, Ariana Bermudez, Joon-Young Lee, Hao LiICCV 2023 · 被引用 69 次
- Multi-Task Learning with Knowledge Distillation for Dense PredictionYangyang Xu, Yibo Yang, Lefei ZhangICCV 2023 · 被引用 18 次
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 被引用 4 次
- Structure Matters: Revisiting Boundary Refinement in Video Object SegmentationGuanyi Qin, Ziyue Wang, Daiyun Shen, Haofeng Liu 等ICCV 2025 · 被引用 1 次
- Token Contrast for Weakly-Supervised Semantic SegmentationLixiang Ru, Heliang Zheng, Yibing Zhan, Bo DuCVPR 2023
它引用的顶会 Paper18
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 被引用 429 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 被引用 398 次
- Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region RefinementYongqing Liang, Xin Li, Navid H. Jafari, Jim ChenNeurIPS 2020 · 被引用 192 次
相关 Paper
- Joint Inductive and Transductive Learning for Video Object SegmentationYunyao Mao, Ning Wang, Wengang Zhou, Houqiang LiICCV 2021 · 被引用 111 次
- Siamese Network with Interactive Transformer for Video Object SegmentationMeng Lan, Jing Zhang, Fengxiang He, Lefei ZhangAAAI 2022 · 被引用 41 次
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li 等CVPR 2023
- Learning Position and Target Consistency for Memory-Based Video Object SegmentationLi Hu, Peng Zhang, Bang Zhang, Pan Pan 等CVPR 2021
- Beyond Pixel and Object: Part Feature as Reference for Few-Shot Video Object SegmentationNaisong Luo, Guoxin Xiong, Tianzhu ZhangAAAI 2025
