Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a Sentence
Weidong Chen, Guorong Li, Xinfeng Zhang, Hongyang Yu, Shuhui Wang, Qingming Huang
摘要
In this paper, we address the problem that selectively segments the actor and its action in the video clip given the sentence description. The main challenge is to match the local semantic features of the video with the heterogeneous textual features. A widely used language processing method in previous works is to leverage bi-LSTM and self-attention, which fixed the attention of the sentence and neglected the personality of the video, leading the attention of the sentence mismatch the most discriminative feature of the video. The proposed algorithm in this paper allows the sentence to learn the most discriminative features of the video, remarkably improving the accuracy of matching and segmentation. Specifically, we propose a cascade cross-modal attention to leverage two perspectives visual features to attend language from coarse to fine to generate the discriminative vision-aware language features. Moreover, equipping our framework with a contrastive learning method and a designed hard negative mining strategy benefits our proposed network from identifying the positive sample from numbers of negatives, and further improving the performance. To demonstrate the effectiveness of our approach, we conduct experiments on two datasets: A2D Sentences and J-HMDB Sentences. Experimental results show that our method significantly improves the performance over recent state-of-the-art methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Multi-Attention Network for Compressed Video Referring Object SegmentationWeidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han 等ACM MM 2022 · 被引用 49 次
- Dual-path Collaborative Generation Network for Emotional Video CaptioningCheng Ye, Weidong Chen, Jingyu Li, Lei Zhang 等ACM MM 2024 · 被引用 15 次
相关 Paper
- Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language QueryHao Wang, Cheng Deng, Junchi Yan, Dacheng TaoICCV 2019 · 被引用 89 次
- Context Modulated Dynamic Networks for Actor and Action Video Segmentation with Language QueriesHao Wang, Cheng Deng, Fan Ma, Yi YangAAAI 2020 · 被引用 58 次
- Visual-Textual Capsule Routing for Text-Based Video SegmentationBruce McIntosh, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahCVPR 2020
- Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor SegmentationTianrui Hui, Shaofei Huang, Si Liu, Zihan Ding 等CVPR 2021
- Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment RetrievalZheng Wang, Jingjing Chen, Yu-Gang JiangACM MM 2021 · 被引用 74 次
