Multimodal Event Causality Reasoning with Scene Graph Enhanced Interaction Network
Jintao Liu, Kaiwen Wei, Chenglong Liu
Abstract
Multimodal event causality reasoning aims to recognize the causal relations based on the given events and accompanying image pairs, requiring the model to have a comprehensive grasp of visual and textual information. However, existing studies fail to effectively model the relations of the objects within the image and capture the object interactions across the image pair, resulting in an insufficient understanding of visual information by the model. To address these issues, we propose a Scene Graph Enhanced Interaction Network (SEIN) in this paper, which can leverage the interactions of the generated scene graph for multimodal event causality reasoning. Specifically, the proposed method adopts a graph convolutional network to model the objects and their relations derived from the scene graph structure, empowering the model to exploit the rich structural and semantic information in the image adequately. To capture the object interactions between the two images, we design an optimal transport-based alignment strategy to match the objects across the images, which could help the model recognize changes in visual information and facilitate causality reasoning. In addition, we introduce a cross-modal fusion module to combine textual and visual features for causality prediction. Experimental results indicate that the proposed SEIN outperforms state-of-the-art methods on the Vis-Causal dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- CLIP-Event: Connecting Text and Images with Event StructuresManling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou et al.CVPR 2022 · 103 citations
- Context-aware Scene Graph Generation with Seq2Seq TransformersYichao Lu, Himanshu Rai, Jason Chang, Boris Knyazev et al.ICCV 2021 · 93 citations
- Data Efficient Language-Supervised Zero-Shot Recognition with Optimal Transport DistillationBichen Wu, Ruizhe Cheng, Peizhao Zhang, Tianren Gao et al.ICLR 2022 · 57 citations
Related papers
- SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense ReasoningZhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian et al.AAAI 2022 · 40 citations
- Two Heads are Better Than One: Hypergraph-Enhanced Graph Reasoning for Visual Event RatiocinationWenbo Zheng, Lan Yan, Chao Gou, Fei-Yue WangICML 2021 · 13 citations
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang et al.AAAI 2020 · 75 citations
- Scene Graph-Grounded Image GenerationFuyun Wang, Tong Zhang, Yuanzhi Wang, Xiaoya Zhang et al.AAAI 2025 · 1 citation
- Visual Semantics Allow for Textual Reasoning Better in Scene Text RecognitionYue He, Chen Chen, Jing Zhang, Juhua Liu et al.AAAI 2022 · 62 citations
