Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual Clues
Qingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng, Shoujie Tong, Haoran Meng, Lin Xu, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu, Zhifang Sui
Abstract
It is a common practice for recent works in vision language cross-modal reasoning to adopt a binary or multi-choice classification formulation taking as input a set of source image(s) and textual query. In this work, we take a sober look at such an "unconditional" formulation in the sense that no prior knowledge is specified with respect to the source image(s). Inspired by the designs of both visual commonsense reasoning and natural language inference tasks, we propose a new task termed "Premise-based Multi-modal Reasoning" (PMR) where a textual premise is the background presumption on each source image. The PMR dataset contains 15,360 manually annotated samples which are created by a multi-phase crowd-sourcing process. With selected high-quality movie screenshots and human-curated premise templates from 6 predefined categories, we ask crowd-source workers to write one true hypothesis and three distractors (4 choices) given the premise and image through a cross-check procedure. Besides, we generate adversarial samples to alleviate the annotation artifacts and double the size of PMR. We benchmark various state-of-theart (pretrained) multi-modal inference models on PMR and conduct comprehensive experimental analyses to showcase the utility of our dataset. * Equal contribution. 1 The dataset and baseline models can be found in https: //2030nlp.github.io/PMR/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc05a709-e24e-4cf9-8caf-f4a0cafd8041Cited by top-tier papers4
- Reasoning with Language Model Prompting: A SurveyShuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen et al.ACL 2023 · 124 citations
- Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and FutureZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu et al.ACL 2024 · 36 citations
- A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesYunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding et al.ACL 2023 · 10 citations
- Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning DistractorJiali Chen, Xusen Hei, Yuqi Xue, Yuancheng Wei et al.ACM MM 2024 · 3 citations
Builds on3
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- What is More Likely to Happen Next? Video-and-Language Future Event PredictionJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalEMNLP 2020 · 45 citations
- Learning Relation Alignment for Calibrated Cross-modal RetrievalShuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men et al.ACL 2021
Related papers
- R1-Onevision: Advancing Generalized Multimodal Reasoning Through Cross-Modal FormalizationYi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang et al.ICCV 2025 · 21 citations
- Violin: A Large-Scale Dataset for Video-and-Language InferenceJingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan et al.CVPR 2020
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop ReasoningJunyoung Sung, Seungwoo Lyu, Minjun Kim, Sumin An et al.CVPR 2026 · 2 citations
- Zero-shot Multimodal Document Retrieval via Cross-modal Question GenerationYejin Choi, Jae-Woo Park, Janghan Yoon, Saejin Kim et al.EMNLP 2025
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long et al.ICLR 2026 · 8 citations
