Multi-Level Representation Learning with Semantic Alignment for Referring Video Object Segmentation
Dongming Wu, Xingping Dong, Ling Shao, Jianbing Shen
摘要
Referring video object segmentation (RVOS) is a challenging language-guided video grounding task, which requires comprehensively understanding the semantic information of both video content and language queries for object prediction. However, existing methods adopt multi-modal fusion at a frame-based spatial granularity. The limitation of visual representation is prone to causing vision-language mismatching and producing poor segmentation results. To address this, we propose a novel multi-level representation learning approach, which explores the inherent structure of the video content to provide a set of discriminative visual embedding, enabling more effective vision-language semantic alignment. Specifically, we embed different visual cues in terms of visual granularity, including multi-frame long-temporal information at video level, intra-frame spatial semantics at frame level, and enhanced object-aware feature prior at object level. With the powerful multi-level visual embedding and carefully-designed dynamic alignment, our model can generate a robust representation for accurate video object segmentation. Extensive experiments on Refer-DAVIS <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">17</inf> and Refer-YouTube-VOS demonstrate that our model achieves superior performance both in segmentation accuracy and inference speed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang 等ICCV 2023 · 被引用 267 次
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsHenghui Ding, Chang Liu, Shuting He, Xudong Jiang 等ICCV 2023 · 被引用 242 次
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosZechen Bai, Tong He, Haiyang Mei, Pichao Wang 等NeurIPS 2024 · 被引用 147 次
- OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationDongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang 等ICCV 2023 · 被引用 82 次
- Spectrum-guided Multi-granularity Referring Video Object SegmentationBo Miao, Mohammed Bennamoun, Yongsheng Gao, Ajmal MianICCV 2023 · 被引用 75 次
它引用的顶会 Paper24
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 被引用 359 次
- Beyond Human Parts: Dual Part-Aligned Representations for Person Re-IdentificationJianyuan Guo, Yuhui Yuan, Lang Huang, Chao Zhang 等ICCV 2019 · 被引用 201 次
相关 Paper
- ReferDINO: Referring Video Object Segmentation with Visual Grounding FoundationsTianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang 等ICCV 2025 · 被引用 7 次
- DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object SegmentationWenxuan Cheng, Ming Dai, Huimin Lu, Wankou YangCVPR 2026
- SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationZhuoyan Luo, Yicheng Xiao, Yong Liu, Shuyan Li 等NeurIPS 2023 · 被引用 89 次
- HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object SegmentationMingfei Han, Yali Wang, Zhihui Li, Lina Yao 等ICCV 2023 · 被引用 42 次
- Tracking-forced Referring Video Object SegmentationRuxue Yan, Wenya Guo, Xubo Liu, Xumeng Liu 等ACM MM 2024 · 被引用 3 次
