Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event Extraction
Jiaqi Li, Chuanyi Zhang, Miaozeng Du, Dehai Min, Yongrui Chen, Guilin Qi
Abstract
Text-video based multimodal event extraction refers to identifying event information from the given text-video pairs. Existing methods predominantly utilize video appearance features (VAF) and text sequence features (TSF) as input information. Some of them employ contrastive learning to align VAF with the event types extracted from TSF. However, they disregard the motion representations in videos and the optimization of contrastive objective could be misguided by the background noise from RGB frames. We observe that the same event triggers correspond to similar motion trajectories, which are hardly affected by the background noise. Moviated by this, we propose a Three Stream Multimodal Event Extraction framework (TSEE) that simultaneously utilizes the features of text sequence and video appearance, as well as the motion representations to enhance the event extraction capacity. Firstly, we extract the optical flow features (OFF) as motion representations from videos to incorporate with VAF and TSF. Then we introduce a Multi-level Event Contrastive Learning module to align the embedding space between OFF and event triggers, as well as between event triggers and types. Finally, a Dual Querying Text module is proposed to enhance the interaction between modalities. Experimental results show that TSEE outperforms the state-of-the-art methods, which demonstrates its superiority.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Revisiting Classical Chinese Event Extraction with Ancient Literature InformationXiaoyi Bao, Zhongqing Wang, Jinghang Gu, Chu-Ren HuangACL 2025
- Evaluation Pitfalls and Challenges in Multimedia Event ExtractionPhilipp Seeberger, Steffen Freisinger, Tobias Bocklet, Korbinian RiedhammerACL 2026
Builds on19
- Supervised Contrastive Learning for Pre-trained Language Model Fine-tuningBeliz Gunel, Jingfei Du, Alexis Conneau, Veselin StoyanovICLR 2021 · 595 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Event Extraction as Machine Reading ComprehensionJian Liu, Yubo Chen, Kang Liu, Wei Bi et al.EMNLP 2020 · 300 citations
- Adversarial Self-Supervised Contrastive LearningMinseon Kim, Jihoon Tack, Sung Ju HwangNeurIPS 2020 · 294 citations
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot LearnersNingyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng et al.ICLR 2022 · 205 citations
Related papers
- Sequence-Event Semantic Consistent Learning for Text-to-Motion RetrievalHaoyu Shi, Huaiwen ZhangACM MM 2025
- Multimedia Event Extraction From News With a Unified Contrastive Learning FrameworkJian Liu, Yufeng Chen, Jinan XuACM MM 2022 · 12 citations
- ATM: Action Temporality Modeling for Video Question AnsweringJunwen Chen, Jie Zhu, Yu KongACM MM 2023 · 1 citation
- Joint Searching and Grounding: Multi-Granularity Video Content RetrievalZhiguo Chen, Xun Jiang, Xing Xu, Zuo Cao et al.ACM MM 2023 · 31 citations
- Multi-event Video-Text RetrievalGengyuan Zhang, Jisen Ren, Jindong Gu, Volker TrespICCV 2023 · 19 citations
