Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation
Shuting He, Henghui Ding
摘要
Referring video segmentation relies on natural language expressions to identify and segment objects, often emphasizing motion clues. Previous works treat a sentence as a whole and directly perform identification at the videolevel, mixing up static image-level cues with temporal motion cues. However, image-level features cannot well comprehend motion cues in sentences, and static cues are not crucial for temporal perception. In fact, static cues can sometimes interfere with temporal perception by overshadowing motion cues. In this work, we propose to decouple video-level referring expression understanding into static and motion perception, with a specific emphasis on enhancing temporal comprehension. Firstly, we introduce an expression-decoupling module to make static cues and motion cues perform their distinct role, alleviating the issue of sentence embeddings overlooking motion cues. Secondly, we propose a hierarchical motion perception module to capture temporal information effectively across varying timescales. Furthermore, we employ contrastive learning to distinguish the motions of visually similar objects. These contributions yield state-of-the-art performance across five datasets, including a remarkable 9.2% J &F improvement on the challenging MeViS dataset. Code is available at https://github.com/heshuting555/DsHmp .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosZechen Bai, Tong He, Haiyang Mei, Pichao Wang 等NeurIPS 2024 · 被引用 147 次
- Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object SegmentationShaofei Huang, Rui Ling, Hongyu Li, Tianrui Hui 等AAAI 2025 · 被引用 24 次
- RefMask3D: Language-Guided Transformer for 3D Referring SegmentationShuting He, Henghui DingACM MM 2024 · 被引用 12 次
- ReferevErything: Towards Segmenting Everything we can Speak of in VideosAnurag Bagchi, Zhipeng Bao, Yu-Xiong Wang, Pavel Tokmakov 等ICCV 2025 · 被引用 11 次
- Deforming Videos to Masks: Flow Matching for Referring Video SegmentationZanyi Wang, Dengyang Jiang, Liuzhuozheng Li, Sizhe Dang 等ICLR 2026 · 被引用 10 次
它引用的顶会 Paper34
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel 等NeurIPS 2020 · 被引用 805 次
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 被引用 359 次
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao 等CVPR 2022 · 被引用 337 次
相关 Paper
- Decoupled Motion Expression Video SegmentationHao Fang, Runmin Cong, Xiankai Lu, Xiaofei Zhou 等CVPR 2025
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsHenghui Ding, Chang Liu, Shuting He, Xudong Jiang 等ICCV 2023 · 被引用 242 次
- Video Moment Retrieval with Hierarchical Contrastive LearningBolin Zhang, Chao Yang, Bin Jiang, Xiaokang ZhouACM MM 2022 · 被引用 21 次
- Temporal Collection and Distribution for Referring Video Object SegmentationJiajin Tang, Ge Zheng, Sibei YangICCV 2023 · 被引用 44 次
- DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object SegmentationWenxuan Cheng, Ming Dai, Huimin Lu, Wankou YangCVPR 2026
