Visual Abductive Reasoning
Chen Liang, Wenguan Wang, Tianfei Zhou, Yi Yang
Abstract
Abductive reasoning seeks the likeliest possible explanation for partial observations. Although abduction is frequently employed in human daily reasoning, it is rarely explored in computer vision literature. In this paper, we propose a new task and dataset, Visual Abductive Reasoning (VAR), for examining abductive reasoning ability of machine intelligence in everyday visual situations. Given an incomplete set of visual events, AI systems are required to not only describe what is observed, but also infer the hypothesis that can best explain the visual premise. Based on our large-scale VAR dataset, we devise a strong baseline model, REASONER (causal-and-cascaded reasoning Transformer). First, to capture the causal structure of the observations, a contextualized directional position embedding strategy is adopted in the encoder, that yields discriminative representations for the premise and hypothesis. Then, multiple decoders are cascaded to generate and progressively refine the premise and hypothesis sentences. The prediction scores of the sentences are used to guide cross-sentence information flow in the cascaded reasoning procedure. Our VAR benchmarking results show that REASONER surpasses many famous video-language models, while still being far behind human performance. This work is expected to foster future efforts in the reasoning-beyond-observation paradigm. * Part of this work was done when Chen Liang was an intern at Baidu.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Decoupling Features in Hierarchical Propagation for Video Object SegmentationZongxin Yang, Yi YangNeurIPS 2022 · 243 citations
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 97 citations
- MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningTieyuan Chen, Huabin Liu, Tianyao He, Yihang Chen et al.NeurIPS 2024 · 36 citations
- DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only TrainingWei Li, Linchao Zhu, Longyin Wen, Yi YangICLR 2023 · 24 citations
- Active Reasoning in an Open-World EnvironmentManjie Xu, Guangyuan Jiang, Wei Liang, Chi Zhang et al.NeurIPS 2023 · 18 citations
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu et al.ICCV 2021 · 427 citations
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
Related papers
- Multi-modal Action Chain Abductive ReasoningMengze Li, Tianbao Wang, Jiahe Xu, Kairong Han et al.ACL 2023 · 11 citations
- AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMsBoyu Chang, Qi Wang, Xi Guo, Zhixiong Nan et al.AAAI 2026 · 1 citation
- Probabilistic Distillation Transformer: Modelling Uncertainties for Visual Abductive ReasoningWanru Xu, Zhenjiang Miao, Yi Tian, Yigang Cen et al.ACM MM 2024
- Transformation Driven Visual ReasoningXin Hong, Yanyan Lan, Liang Pang, Jiafeng Guo et al.CVPR 2021
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
