Visual-Textual Capsule Routing for Text-Based Video Segmentation
Bruce McIntosh, Kevin Duarte, Yogesh S. Rawat, Mubarak Shah
Abstract
Joint understanding of vision and natural language is a challenging problem with a wide range of applications in artificial intelligence. In this work, we focus on integration of video and text for the task of actor and action video segmentation from a sentence. We propose a capsule-based approach which performs pixel-level localization based on a natural language query describing the actor of interest. We encode both the video and textual input in the form of capsules, which provide a more effective representation in comparison with standard convolution based features. Our novel visual-textual routing mechanism allows for the fusion of video and text capsules to successfully localize the actor and action. The existing works on actor-action localization are mainly focused on localization in a single frame instead of the full video. Different from existing works, we propose to perform the localization on all frames of the video. To validate the potential of the proposed network for actor and action video localization, we extend an existing actor-action dataset (A2D) with annotations for all the frames. The experimental evaluation demonstrates the effectiveness of our capsule network for text selective actor and action localization in videos. The proposed method also improves upon the performance of the existing stateof-the art works on single frame-based localization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5ba1887-0e29-4433-ba5c-b10bd582d103Cited by top-tier papers13
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 242 citations
- End-to-End Referring Video Object Segmentation with Multimodal TransformersAdam Botach, Evgenii Zheltonozhskii, Chaim BaskinCVPR 2022 · 150 citations
- SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationZhuoyan Luo, Yicheng Xiao, Yong Liu, Shuyan Li et al.NeurIPS 2023 · 89 citations
- Language-Bridged Spatial-Temporal Interaction for Referring Video Object SegmentationZihan Ding, Tianrui Hui, Junshi Huang, Xiaoming Wei et al.CVPR 2022 · 62 citations
- Multi-Level Representation Learning with Semantic Alignment for Referring Video Object SegmentationDongming Wu, Xingping Dong, Ling Shao, Jianbing ShenCVPR 2022 · 55 citations
Related papers
- Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language QueryHao Wang, Cheng Deng, Junchi Yan, Dacheng TaoICCV 2019 · 89 citations
- Capsule-based Object Tracking with Natural Language SpecificationDing Ma, Xiangqian WuACM MM 2021 · 25 citations
- Context Modulated Dynamic Networks for Actor and Action Video Segmentation with Language QueriesHao Wang, Cheng Deng, Fan Ma, Yi YangAAAI 2020 · 58 citations
- Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a SentenceWeidong Chen, Guorong Li, Xinfeng Zhang, Hongyang Yu et al.ACM MM 2021 · 16 citations
- Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor SegmentationTianrui Hui, Shaofei Huang, Si Liu, Zihan Ding et al.CVPR 2021
