Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited Samples
Guanghui Li, Mingqi Gao, Heng Liu, Xiantong Zhen, Feng Zheng
摘要
Referring video object segmentation (RVOS), as a supervised learning task, relies on sufficient annotated data for a given scene. However, in more realistic scenarios, only minimal annotations are available for a new scene, which poses significant challenges to existing RVOS methods. With this in mind, we propose a simple yet effective model with a newly designed cross-modal affinity (CMA) module based on a Transformer architecture. The CMA module builds multimodal affinity with a few samples, thus quickly learning new semantic information, and enabling the model to adapt to different scenarios. Since the proposed method targets limited samples for new scenes, we generalize the problem as - few-shot referring video object segmentation (FS-RVOS). To foster research in this direction, we build up a new FS-RVOS benchmark based on currently available datasets. The benchmark covers a wide range and includes multiple situations, which can maximally simulate real-world scenarios. Extensive experiments show that our model adapts well to different scenarios with only a few samples, reaching state-of-the-art performance on the benchmark. On Mini-Ref-YouTube-VOS, our model achieves an average performance of 53.1 and 54.8 , which are 10% better than the baselines. Furthermore, we show impressive results of 77.7 and 74.8 on Mini-Ref-SAIL-VOS, which are significantly better than the baselines. Code is publicly available at https://github.com/hengliusky/Few_shot_RVOS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou 等ICCV 2019 · 被引用 1,404 次
- Feature Weighting and Boosting for Few-Shot SegmentationKhoi Nguyen, Sinisa TodorovicICCV 2019 · 被引用 402 次
- Pyramid Graph Networks With Connection Attentions for Region-Based One-Shot Semantic SegmentationChi Zhang, Guosheng Lin, Fayao Liu, Jiushuang Guo 等ICCV 2019 · 被引用 351 次
相关 Paper
- Segment Anything Across Shots: A Method and BenchmarkHengrui Hu, Kaining Ying, Henghui DingAAAI 2026 · 被引用 1 次
- End-to-End Referring Video Object Segmentation with Multimodal TransformersAdam Botach, Evgenii Zheltonozhskii, Chaim BaskinCVPR 2022 · 被引用 150 次
- OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationDongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang 等ICCV 2023 · 被引用 82 次
- Video State-Changing Object SegmentationJiangwei Yu, Xiang Li, Xinran Zhao, Hongming Zhang 等ICCV 2023 · 被引用 16 次
- Two-shot Video Object SegmentationKun Yan, Xiao Li, Fangyun Wei, Jinglu Wang 等CVPR 2023
