Explainable Video Entailment with Grounded Visual Evidence
Junwen Chen, Yu Kong Golisano
摘要
Video entailment aims at determining if a hypothesis textual statement is entailed or contradicted by a premise video. The main challenge of video entailment is that it requires fine-grained reasoning to understand the complex and long story-based videos. To this end, we propose to incorporate visual grounding to the entailment by explicitly linking the entities described in the statement to the evidence in the video. If the entities are grounded in the video, we enhance the entailment judgment by focusing on the frames where the entities occur. Besides, in the entailment dataset, the entailed/contradictory (also named as real/fake) statements are formed in pairs with subtle discrepancy, which allows an add-on explanation module to predict which words or phrases make the statement contradictory to the video and regularize the training of the entailment judgment. Experimental results demonstrate that our approach outperforms the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu 等NeurIPS 2022 · 被引用 91 次
- i-Code: An Integrative and Composable Multimodal Learning FrameworkZiyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant 等AAAI 2023 · 被引用 53 次
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 被引用 44 次
- TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video ReasoningKate Sanders, Nathaniel Weir, Benjamin Van DurmeEMNLP 2024 · 被引用 3 次
- ART: rule bAsed futuRe-inference deducTionMengze Li, Tianqi Zhao, Jionghao Bai, Baoyi He 等EMNLP 2023 · 被引用 2 次
它引用的顶会 Paper6
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 被引用 173 次
- Activity-driven Weakly-Supervised Spatio-Temporal Grounding from Untrimmed VideosJunwen Chen, Wentao Bao, Yu KongACM MM 2020 · 被引用 19 次
- Violin: A Large-Scale Dataset for Video-and-Language InferenceJingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan 等CVPR 2020
- Modality Shifting Attention Network for Multi-Modal Video Question AnsweringJunyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim 等CVPR 2020
相关 Paper
- Comprehensive Visual Grounding for Video DescriptionWenhui Jiang, Yibo Cheng, Linxin Liu, Yuming Fang 等AAAI 2024 · 被引用 5 次
- Video Entailment via Reaching a Structure-Aware Cross-modal ConsensusXuan Yao, Junyu Gao, Mengyuan Chen, Changsheng XuACM MM 2023 · 被引用 4 次
- Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Wei Ji 等CVPR 2022 · 被引用 108 次
- Scene-text Oriented Visual Entailment: Task, Dataset and SolutionNan Li, Pijian Li, Dongsheng Xu, Wenye Zhao 等ACM MM 2023 · 被引用 2 次
- CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video GroundingYaoyuan Liang, Xiao Liang, Yansong Tang, Zhao Yang 等AAAI 2024 · 被引用 3 次
