Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a Sentence
Weidong Chen, Guorong Li, Xinfeng Zhang, Hongyang Yu, Shuhui Wang, Qingming Huang
Abstract
In this paper, we address the problem that selectively segments the actor and its action in the video clip given the sentence description. The main challenge is to match the local semantic features of the video with the heterogeneous textual features. A widely used language processing method in previous works is to leverage bi-LSTM and self-attention, which fixed the attention of the sentence and neglected the personality of the video, leading the attention of the sentence mismatch the most discriminative feature of the video. The proposed algorithm in this paper allows the sentence to learn the most discriminative features of the video, remarkably improving the accuracy of matching and segmentation. Specifically, we propose a cascade cross-modal attention to leverage two perspectives visual features to attend language from coarse to fine to generate the discriminative vision-aware language features. Moreover, equipping our framework with a contrastive learning method and a designed hard negative mining strategy benefits our proposed network from identifying the positive sample from numbers of negatives, and further improving the performance. To demonstrate the effectiveness of our approach, we conduct experiments on two datasets: A2D Sentences and J-HMDB Sentences. Experimental results show that our method significantly improves the performance over recent state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 537fc12a-32f8-4cd7-af99-1f4f15b4f61dCited by top-tier papers2
- Multi-Attention Network for Compressed Video Referring Object SegmentationWeidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han et al.ACM MM 2022 · 49 citations
- Dual-path Collaborative Generation Network for Emotional Video CaptioningCheng Ye, Weidong Chen, Jingyu Li, Lei Zhang et al.ACM MM 2024 · 15 citations
Related papers
- Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language QueryHao Wang, Cheng Deng, Junchi Yan, Dacheng TaoICCV 2019 · 89 citations
- Context Modulated Dynamic Networks for Actor and Action Video Segmentation with Language QueriesHao Wang, Cheng Deng, Fan Ma, Yi YangAAAI 2020 · 58 citations
- Visual-Textual Capsule Routing for Text-Based Video SegmentationBruce McIntosh, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahCVPR 2020
- Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor SegmentationTianrui Hui, Shaofei Huang, Si Liu, Zihan Ding et al.CVPR 2021
- Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment RetrievalZheng Wang, Jingjing Chen, Yu-Gang JiangACM MM 2021 · 74 citations
