Modeling Motion with Multi-Modal Features for Text-Based Video Segmentation
Wangbo Zhao, Kai Wang, Xiangxiang Chu, Fuzhao Xue, Xinchao Wang, Yang You
摘要
Text-based video segmentation aims to segment the target object in a video based on a describing sentence. Incorporating motion information from optical flow maps with appearance and linguistic modalities is crucial yet has been largely ignored by previous work. In this paper, we design a method to fuse and align appearance, motion, and linguistic features to achieve accurate segmentation. Specifically, we propose a multi-modal video transformer, which can fuse and aggregate multi-modal and temporal features between frames. Furthermore, we design a language-guided feature fusion module to progressively fuse appearance and motion features in each feature level with guidance from linguistic features. Finally, a multi-modal alignment loss is proposed to alleviate the semantic gap between features from different modalities. Extensive experiments on A2D Sentences and J-HMDB Sentences verify the performance and the generalization ability of our method compared to the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsHenghui Ding, Chang Liu, Shuting He, Xudong Jiang 等ICCV 2023 · 被引用 242 次
- OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationDongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang 等ICCV 2023 · 被引用 82 次
- Spectrum-guided Multi-granularity Referring Video Object SegmentationBo Miao, Mohammed Bennamoun, Yongsheng Gao, Ajmal MianICCV 2023 · 被引用 75 次
- Temporal Collection and Distribution for Referring Video Object SegmentationJiajin Tang, Ge Zheng, Sibei YangICCV 2023 · 被引用 44 次
- HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object SegmentationMingfei Han, Yali Wang, Zhihui Li, Lina Yao 等ICCV 2023 · 被引用 42 次
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- End-to-end Multi-modal Video Temporal GroundingYi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan YangNeurIPS 2021 · 被引用 68 次
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu 等CVPR 2022 · 被引用 146 次
- Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationShilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen 等AAAI 2024 · 被引用 67 次
- M3L: Language-based Video Editing via Multi-Modal Multi-Level TransformersTsu-Jui Fu, Xin Eric Wang, Scott T. Grafton, Miguel P. Eckstein 等CVPR 2022 · 被引用 13 次
- End-to-End Referring Video Object Segmentation with Multimodal TransformersAdam Botach, Evgenii Zheltonozhskii, Chaim BaskinCVPR 2022 · 被引用 150 次
