Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning
Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, Alan L. Yuille
Abstract
Visual Question Answering (VQA) models often perform poorly on out-of-distribution data and struggle on domain generalization. Due to the multi-modal nature of this task, multiple factors of variation are intertwined, making generalization difficult to analyze. This motivates us to introduce a virtual benchmark, Super-CLEVR, where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently. Four factors are considered: visual complexity, question redundancy, concept distribution and concept compositionality. With controllably generated data, Super-CLEVR enables us to test VQA methods in situations where the test data differs from the training data along each of these axes. We study four existing methods, including two neural symbolic methods NSCL [45] and NSVQA [59], and two non-symbolic methods FiLM [50] and mDETR [29]; and our proposed method, probabilistic NSVQA (P-NSVQA), which extends NSVQA with uncertainty reasoning. P-NSVQA outperforms other methods on three of the four domain shift factors. Our results suggest that disentangling reasoning and perception, combined with probabilistic uncertainty, form a strong VQA model that is more robust to domain shifts. The dataset and code are released at https://github.com/Lizw14/ Super-CLEVR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ab40b3e-02fe-4606-a764-7cc8c0c67e82Cited by top-tier papers50
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HDXiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao et al.NeurIPS 2024 · 193 citations
- Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree SearchHuanjin Yao, Jiaxing Huang, Wenhao Wu, Jingyi Zhang et al.NeurIPS 2025 · 147 citations
- Perception-Aware Policy Optimization for Multimodal ReasoningZhenhailong Wang, Xuehang Guo, Sofia Stoica, Haiyang Xu et al.ICLR 2026 · 104 citations
Builds on17
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Episodic Training for Domain GeneralizationDa Li, Jianshu Zhang, Yongxin Yang, Cong Liu et al.ICCV 2019 · 488 citations
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 391 citations
- Re-distributing Biased Pseudo Labels for Semi-supervised Semantic Segmentation: A Baseline InvestigationRuifei He, Jihan Yang, Xiaojuan QiICCV 2021 · 149 citations
Related papers
- Are Object-Centric Representations Better at Compositional Generalization?Ferdinand Kapl, Amir Mohammad Karimi Mamaghan, Maximilian Seitzer, Karl Johansson et al.ICML 2026
- 3D-Aware Visual Question Answering about Parts, Poses and OcclusionsXingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski et al.NeurIPS 2023 · 27 citations
- Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question AnsweringXingrui Wang, Wufei Ma, Angtian Wang, Shuo Chen et al.ICLR 2025
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma et al.EMNLP 2021 · 18 citations
- Domain-Robust VQA With Diverse Datasets and Methods but No Target LabelsMingda Zhang, Tristan Maidment, Ahmad Diab, Adriana Kovashka et al.CVPR 2021
