Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event Extraction
Jiaqi Li, Chuanyi Zhang, Miaozeng Du, Dehai Min, Yongrui Chen, Guilin Qi
摘要
Text-video based multimodal event extraction refers to identifying event information from the given text-video pairs. Existing methods predominantly utilize video appearance features (VAF) and text sequence features (TSF) as input information. Some of them employ contrastive learning to align VAF with the event types extracted from TSF. However, they disregard the motion representations in videos and the optimization of contrastive objective could be misguided by the background noise from RGB frames. We observe that the same event triggers correspond to similar motion trajectories, which are hardly affected by the background noise. Moviated by this, we propose a Three Stream Multimodal Event Extraction framework (TSEE) that simultaneously utilizes the features of text sequence and video appearance, as well as the motion representations to enhance the event extraction capacity. Firstly, we extract the optical flow features (OFF) as motion representations from videos to incorporate with VAF and TSF. Then we introduce a Multi-level Event Contrastive Learning module to align the embedding space between OFF and event triggers, as well as between event triggers and types. Finally, a Dual Querying Text module is proposed to enhance the interaction between modalities. Experimental results show that TSEE outperforms the state-of-the-art methods, which demonstrates its superiority.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Revisiting Classical Chinese Event Extraction with Ancient Literature InformationXiaoyi Bao, Zhongqing Wang, Jinghang Gu, Chu-Ren HuangACL 2025
- Evaluation Pitfalls and Challenges in Multimedia Event ExtractionPhilipp Seeberger, Steffen Freisinger, Tobias Bocklet, Korbinian RiedhammerACL 2026
它引用的顶会 Paper19
- Supervised Contrastive Learning for Pre-trained Language Model Fine-tuningBeliz Gunel, Jingfei Du, Alexis Conneau, Veselin StoyanovICLR 2021 · 被引用 595 次
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 被引用 522 次
- Event Extraction as Machine Reading ComprehensionJian Liu, Yubo Chen, Kang Liu, Wei Bi 等EMNLP 2020 · 被引用 300 次
- Adversarial Self-Supervised Contrastive LearningMinseon Kim, Jihoon Tack, Sung Ju HwangNeurIPS 2020 · 被引用 294 次
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot LearnersNingyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng 等ICLR 2022 · 被引用 205 次
相关 Paper
- Sequence-Event Semantic Consistent Learning for Text-to-Motion RetrievalHaoyu Shi, Huaiwen ZhangACM MM 2025
- Multimedia Event Extraction From News With a Unified Contrastive Learning FrameworkJian Liu, Yufeng Chen, Jinan XuACM MM 2022 · 被引用 12 次
- ATM: Action Temporality Modeling for Video Question AnsweringJunwen Chen, Jie Zhu, Yu KongACM MM 2023 · 被引用 1 次
- Joint Searching and Grounding: Multi-Granularity Video Content RetrievalZhiguo Chen, Xun Jiang, Xing Xu, Zuo Cao 等ACM MM 2023 · 被引用 31 次
- Multi-event Video-Text RetrievalGengyuan Zhang, Jisen Ren, Jindong Gu, Volker TrespICCV 2023 · 被引用 19 次
