Multi-modal Action Chain Abductive Reasoning
Mengze Li, Tianbao Wang, Jiahe Xu, Kairong Han, Shengyu Zhang, Zhou Zhao, Jiaxu Miao, Wenqiao Zhang, Shiliang Pu, Fei Wu
Abstract
Abductive Reasoning, has long been considered to be at the core ability of humans, which enables us to infer the most plausible explanation of incomplete known phenomena in daily life. However, such critical reasoning capability is rarely investigated for contemporary AI systems under such limited observations. To facilitate this research community, this paper sheds new light on Abductive Reasoning by studying a new vision-language task, Multi-modal Action chain abductive Reasoning (MAR), together with a large-scale Abductive Reasoning dataset: Given an incomplete set of language described events, MAR aims to imagine the most plausible event by spatio-temporal grounding in past video and then infer the hypothesis of subsequent action chain that can best explain the language premise. To solve this task, we propose a strong baseline model that realizes MAR from two perspectives: (i) we first introduce the transformer, which learns to encode the observation to imagine the plausible event with explicitly interpretable event grounding in the video based on the commonsense knowledge recognition ability. (ii) To complete the assumption of a follow-up action chain, we design a novel symbolic module that can complete strict derivation of the progressive action chain layer by layer. We conducted extensive experiments on the proposed dataset, and the experimental study shows that the proposed model significantly outperforms existing video-language models in terms of effectiveness on our newly created MAR dataset. Our dataset is available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66596fec-d9d8-43ef-a306-20ca836c2b9cCited by top-tier papers5
- Intelligent Model Update Strategy for Sequential RecommendationZheqi Lv, Wenqiao Zhang, Zhengyu Chen, Shengyu Zhang et al.WWW 2024 · 53 citations
- Revisiting the Domain Shift and Sample Uncertainty in Multi-source Active Domain TransferWenqiao Zhang, Zheqi LvCVPR 2024 · 17 citations
- Diagrammatization and Abduction to Improve AI Interpretability With Domain-Aligned Explanations for Medical DiagnosisBrian Y. Lim, Joseph P. Cahaly, Chester Y. F. Sng, Adam ChewCHI 2025 · 11 citations
- Quantitatively Measuring and Contrastively Exploring Heterogeneity for Domain GeneralizationYunze Tong, Junkun Yuan, Min Zhang, Didi Zhu et al.KDD 2023 · 7 citations
- AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMsBoyu Chang, Qi Wang, Xi Guo, Zhixiong Nan et al.AAAI 2026 · 1 citation
Builds on17
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji et al.AAAI 2021 · 186 citations
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo et al.NeurIPS 2021 · 90 citations
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear et al.ICML 2020 · 88 citations
Related papers
- Visual Abductive ReasoningChen Liang, Wenguan Wang, Tianfei Zhou, Yi YangCVPR 2022 · 50 citations
- Cross-modal Observation Hypothesis InferenceMengze Li, Kairong Han, Jiahe Xu, Yueying Li et al.ACM MM 2024
- REX: Reasoning-aware and Grounded ExplanationShi Chen, Qi ZhaoCVPR 2022 · 24 citations
- Multimodal Analogical Reasoning over Knowledge GraphsNingyu Zhang, Lei Li, Xiang Chen, Xiaozhuan Liang et al.ICLR 2023 · 10 citations
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
