Explainable Video Entailment with Grounded Visual Evidence
Junwen Chen, Yu Kong Golisano
Abstract
Video entailment aims at determining if a hypothesis textual statement is entailed or contradicted by a premise video. The main challenge of video entailment is that it requires fine-grained reasoning to understand the complex and long story-based videos. To this end, we propose to incorporate visual grounding to the entailment by explicitly linking the entities described in the statement to the evidence in the video. If the entities are grounded in the video, we enhance the entailment judgment by focusing on the frames where the entities occur. Besides, in the entailment dataset, the entailed/contradictory (also named as real/fake) statements are formed in pairs with subtle discrepancy, which allows an add-on explanation module to predict which words or phrases make the statement contradictory to the video and regularize the training of the entailment judgment. Experimental results demonstrate that our approach outperforms the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae2b1530-b3b1-49ab-a2db-74b639526befCited by top-tier papers6
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu et al.NeurIPS 2022 · 91 citations
- i-Code: An Integrative and Composable Multimodal Learning FrameworkZiyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant et al.AAAI 2023 · 53 citations
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 44 citations
- TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video ReasoningKate Sanders, Nathaniel Weir, Benjamin Van DurmeEMNLP 2024 · 3 citations
- ART: rule bAsed futuRe-inference deducTionMengze Li, Tianqi Zhao, Jionghao Bai, Baoyi He et al.EMNLP 2023 · 2 citations
Builds on6
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 173 citations
- Activity-driven Weakly-Supervised Spatio-Temporal Grounding from Untrimmed VideosJunwen Chen, Wentao Bao, Yu KongACM MM 2020 · 19 citations
- Violin: A Large-Scale Dataset for Video-and-Language InferenceJingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan et al.CVPR 2020
- Modality Shifting Attention Network for Multi-Modal Video Question AnsweringJunyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim et al.CVPR 2020
Related papers
- Comprehensive Visual Grounding for Video DescriptionWenhui Jiang, Yibo Cheng, Linxin Liu, Yuming Fang et al.AAAI 2024 · 5 citations
- Video Entailment via Reaching a Structure-Aware Cross-modal ConsensusXuan Yao, Junyu Gao, Mengyuan Chen, Changsheng XuACM MM 2023 · 4 citations
- Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Wei Ji et al.CVPR 2022 · 108 citations
- Scene-text Oriented Visual Entailment: Task, Dataset and SolutionNan Li, Pijian Li, Dongsheng Xu, Wenye Zhao et al.ACM MM 2023 · 2 citations
- CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video GroundingYaoyuan Liang, Xiao Liang, Yansong Tang, Zhao Yang et al.AAAI 2024 · 3 citations
