Visual Abductive Reasoning
Chen Liang, Wenguan Wang, Tianfei Zhou, Yi Yang
摘要
Abductive reasoning seeks the likeliest possible explanation for partial observations. Although abduction is frequently employed in human daily reasoning, it is rarely explored in computer vision literature. In this paper, we propose a new task and dataset, Visual Abductive Reasoning (VAR), for examining abductive reasoning ability of machine intelligence in everyday visual situations. Given an incomplete set of visual events, AI systems are required to not only describe what is observed, but also infer the hypothesis that can best explain the visual premise. Based on our large-scale VAR dataset, we devise a strong baseline model, REASONER (causal-and-cascaded reasoning Transformer). First, to capture the causal structure of the observations, a contextualized directional position embedding strategy is adopted in the encoder, that yields discriminative representations for the premise and hypothesis. Then, multiple decoders are cascaded to generate and progressively refine the premise and hypothesis sentences. The prediction scores of the sentences are used to guide cross-sentence information flow in the cascaded reasoning procedure. Our VAR benchmarking results show that REASONER surpasses many famous video-language models, while still being far behind human performance. This work is expected to foster future efforts in the reasoning-beyond-observation paradigm. * Part of this work was done when Chen Liang was an intern at Baidu.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Decoupling Features in Hierarchical Propagation for Video Object SegmentationZongxin Yang, Yi YangNeurIPS 2022 · 被引用 243 次
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 被引用 97 次
- MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningTieyuan Chen, Huabin Liu, Tianyao He, Yihang Chen 等NeurIPS 2024 · 被引用 36 次
- DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only TrainingWei Li, Linchao Zhu, Longyin Wen, Yi YangICLR 2023 · 被引用 24 次
- Active Reasoning in an Open-World EnvironmentManjie Xu, Guangyuan Jiang, Wei Liang, Chi Zhang 等NeurIPS 2023 · 被引用 18 次
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu 等ICCV 2021 · 被引用 427 次
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 被引用 358 次
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
相关 Paper
- Multi-modal Action Chain Abductive ReasoningMengze Li, Tianbao Wang, Jiahe Xu, Kairong Han 等ACL 2023 · 被引用 11 次
- AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMsBoyu Chang, Qi Wang, Xi Guo, Zhixiong Nan 等AAAI 2026 · 被引用 1 次
- Probabilistic Distillation Transformer: Modelling Uncertainties for Visual Abductive ReasoningWanru Xu, Zhenjiang Miao, Yi Tian, Yigang Cen 等ACM MM 2024
- Transformation Driven Visual ReasoningXin Hong, Yanyan Lan, Liang Pang, Jiafeng Guo 等CVPR 2021
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
