End-to-end Multi-modal Video Temporal Grounding
Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang
摘要
We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as visual features, we propose a multi-modal framework to extract complementary information from videos. Specifically, we adopt RGB images for appearance, optical flow for motion, and depth maps for image structure. While RGB images provide abundant visual cues of certain events, the performance may be affected by background clutters. Therefore, we use optical flow to focus on large motion and depth maps to infer the scene configuration when the action is related to objects recognizable with their shapes. To integrate the three modalities more effectively and enable inter-modal learning, we design a dynamic fusion scheme with transformers to model the interactions between modalities. Furthermore, we apply intra-modal self-supervised learning to enhance feature representations across videos for each modality, which also facilitates multi-modal learning. We conduct extensive experiments on the Charades-STA and ActivityNet Captions datasets, and show that the proposed method performs favorably against state-of-the-art approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao 等NeurIPS 2023 · 被引用 113 次
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon 等ICCV 2023 · 被引用 103 次
- Video Moment Retrieval from Text Queries via Single Frame AnnotationRan Cui, Tianwen Qian, Pai Peng, Elena Daskalaki 等SIGIR 2022 · 被引用 42 次
- Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video GroundingPeijun Bao, Yong Xia, Wenhan Yang, Boon Poh Ng 等AAAI 2024 · 被引用 20 次
- OmniViD: A Generative Framework for Universal Video UnderstandingJunke Wang, Dongdong Chen, Chong Luo, Bo He 等CVPR 2024 · 被引用 18 次
它引用的顶会 Paper9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu 等NeurIPS 2020 · 被引用 561 次
- Learning Cross-Modal Contrastive Features for Video Domain AdaptationDonghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu 等ICCV 2021 · 被引用 88 次
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh 等CVPR 2020
相关 Paper
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
- Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationWangbo Zhao, Kai Wang, Xiangxiang Chu, Fuzhao Xue 等CVPR 2022 · 被引用 23 次
- Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence GroundingJiaming Chen, Weixin Luo, Wei Zhang, Lin MaAAAI 2022 · 被引用 33 次
- Text-Visual Prompting for Efficient 2D Temporal Video GroundingYimeng Zhang, Xin Chen, Jinghan Jia, Sijia Liu 等CVPR 2023
- DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-to-Fine Contrastive RankingLijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl 等CVPR 2023
