Multi-event Video-Text Retrieval
Gengyuan Zhang, Jisen Ren, Jindong Gu, Volker Tresp
摘要
Video-Text Retrieval (VTR) is a crucial multi-modal task in an era of massive video-text data on the Internet. A plethora of work characterized by using a two-stream Vision-Language model architecture that learns a joint representation of video-text pairs has become a prominent approach for the VTR task. However, these models operate under the assumption of bijective video-text correspondences and neglect a more practical scenario where video content usually encompasses multiple events, while texts like user queries or webpage metadata tend to be specific and correspond to single events. This establishes a gap between the previous training objective and real-world applications, leading to the potential performance degradation of earlier models during inference. In this study, we introduce the Multi-event Video-Text Retrieval (MeVTR) task, addressing scenarios in which each video contains multiple different events, as a niche scenario of the conventional Video-Text Retrieval Task. We present a simple model, Me-Retriever, which incorporates key event video representation and a new MeVTR loss for the MeVTR task. Comprehensive experiments show that this straightforward framework outperforms other models in the Video-to-Text and Text-to-Video tasks, effectively establishing a robust baseline for the MeVTR task. We believe this work serves as a strong foundation for future studies. Code is available at https://github.com/gengyuanmax/MeVTR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 等ACM MM 2022 · 被引用 65 次
- SHiNe: Semantic Hierarchy Nexus for Open-Vocabulary Object DetectionMingxuan Liu, Tyler L. Hayes, Elisa Ricci, Gabriela Csurka 等CVPR 2024
- Localizing Events in Videos with Multimodal QueriesGengyuan Zhang, Mang Ling Ada Fok, Jialu Ma, Yan Xia 等CVPR 2025
- TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsXin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen 等CVPR 2025
- Gravitation-Driven Semantic Alignment for Text Video RetrievalYi Yang, Zheng Wang, Xing Xu, Jingkuan Song 等CVPR 2026
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalChengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang 等NeurIPS 2022 · 被引用 52 次
- Joint Searching and Grounding: Multi-Granularity Video Content RetrievalZhiguo Chen, Xun Jiang, Xing Xu, Zuo Cao 等ACM MM 2023 · 被引用 31 次
- Clover: Towards A Unified Video-Language Alignment and Fusion ModelJingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu 等CVPR 2023
- Maskable Retentive Network for Video Moment RetrievalJingjing Hu, Dan Guo, Kun Li, Zhan Si 等ACM MM 2024 · 被引用 7 次
- Relation Triplet Construction for Cross-modal Text-to-Video RetrievalXue Song, Jingjing Chen, Yu-Gang JiangACM MM 2023 · 被引用 8 次
