ART: rule bAsed futuRe-inference deducTion
Mengze Li, Tianqi Zhao, Jionghao Bai, Baoyi He, Jiaxu Miao, Wei Ji, Zheqi Lv, Zhou Zhao, Shengyu Zhang, Wenqiao Zhang, Fei Wu
摘要
Deductive reasoning is a crucial cognitive ability of humanity, allowing us to derive valid conclusions from premises and observations. However, existing works mainly focus on languagebased premises and generally neglect deductive reasoning from visual observations. In this work, we introduce rule bAsed futuReinference deducTion (ART), which aims at deducing the correct future event based on the visual phenomenon (a video) and the rule-based premises, along with an explanation of the reasoning process. To advance this field, we construct a large-scale densely annotated dataset (Video-ART), where the premises, future event candidates, the reasoning process explanation, and auxiliary commonsense knowledge (e.g., actions and appearance) are annotated by annotators. Upon Video-ART, we develop a strong baseline named ARTNet. In essence, guided by commonsense knowledge, ARTNet learns to identify the target video character and perceives its visual clues related to the future event. Then, ARTNet rigorously applies the given premises to conduct reasoning from the identified information to future events, through a non-parametric rule reasoning network and a reasoning-path review module. Empirical studies validate the rationality of ARTNet in deductive reasoning upon visual observations and the effectiveness over existing works.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan 等EMNLP 2020 · 被引用 387 次
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu 等NeurIPS 2022 · 被引用 91 次
- De-Biased Court's View Generation with CausalityYiquan Wu, Kun Kuang, Yating Zhang, Xiaozhong Liu 等EMNLP 2020 · 被引用 71 次
- Consensus Graph Representation Learning for Better Grounded Image CaptioningWenqiao Zhang, Haochen Shi, Siliang Tang, Jun Xiao 等AAAI 2021 · 被引用 63 次
- Large-scale Video Panoptic Segmentation in the Wild: A BenchmarkJiaxu Miao, Xiaohan Wang, Yu Wu, Wei Li 等CVPR 2022 · 被引用 58 次
相关 Paper
- Multi-modal Action Chain Abductive ReasoningMengze Li, Tianbao Wang, Jiahe Xu, Kairong Han 等ACL 2023 · 被引用 11 次
- Cross-modal Observation Hypothesis InferenceMengze Li, Kairong Han, Jiahe Xu, Yueying Li 等ACM MM 2024
- Visual Abductive ReasoningChen Liang, Wenguan Wang, Tianfei Zhou, Yi YangCVPR 2022 · 被引用 50 次
- Violin: A Large-Scale Dataset for Video-and-Language InferenceJingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan 等CVPR 2020
- What is More Likely to Happen Next? Video-and-Language Future Event PredictionJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalEMNLP 2020 · 被引用 45 次
