Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
Shantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston Tan
Abstract
Complex visual reasoning and question answering (VQA) is a challenging task that requires compositional multi-step processing and higher-level reasoning capabilities beyond the immediate recognition and localization of objects and events. Here, we introduce a fully neural Iterative and Parallel Reasoning Mechanism (IPRM) that combines two distinct forms of computation -- iterative and parallel -- to better address complex VQA scenarios. Specifically, IPRM's"iterative"computation facilitates compositional step-by-step reasoning for scenarios wherein individual operations need to be computed, stored, and recalled dynamically (e.g. when computing the query"determine the color of pen to the left of the child in red t-shirt sitting at the white table"). Meanwhile, its"parallel"computation allows for the simultaneous exploration of different reasoning paths and benefits more robust and efficient execution of operations that are mutually independent (e.g. when counting individual colors for the query:"determine the maximum occurring color amongst all t-shirts"). We design IPRM as a lightweight and fully-differentiable neural module that can be conveniently applied to both transformer and non-transformer vision-language backbones. It notably outperforms prior task-specific methods and transformer-based attention modules across various image and video VQA benchmarks testing distinct complex reasoning capabilities such as compositional spatiotemporal reasoning (AGQA), situational reasoning (STAR), multi-hop reasoning generalization (CLEVR-Humans) and causal event linking (CLEVRER-Humans). Further, IPRM's internal computations can be visualized across reasoning steps, aiding interpretability and diagnosis of its errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cf9db25-8f6e-4a10-869b-1ed7829fe73dCited by top-tier papers3
- Imore: Implicit Program-Guided Reasoning for Human Motion QAChen Li, Chinthani Sugandhika, Ee Yeo Keat, Eric P. Xing et al.ICCV 2025 · 7 citations
- PKR-QA: A Benchmark for Procedural Knowledge Reasoning with Knowledge Module LearningThanh-Son Nguyen, Hong Yang, Tzeh Yuan Neoh, Hao Zhang et al.AAAI 2026 · 1 citation
- Towards Robust Uncertainty Calibration for Composed Image RetrievalYifan Wang, Wuliang Huang, Yufan Wen, Shunning Liu et al.NeurIPS 2025
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- MIST : Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question AnsweringDifei Gao, Luowei Zhou, Lei Ji, Linchao Zhu et al.CVPR 2023
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question AnsweringYuanhao Zou, Shengji Jin, Andong Deng, Youpeng Zhao et al.ICLR 2026
- Focal and Composed Vision-semantic Modeling for Visual Question AnsweringYudong Han, Yangyang Guo, Jianhua Yin, Meng Liu et al.ACM MM 2021 · 14 citations
- Dynamic Spatio-Temporal Modular Network for Video Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 14 citations
- Query and Attention Augmentation for Knowledge-Based Explainable ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2022 · 14 citations
