CARETS: A Consistency And Robustness Evaluative Test Suite for VQA
Carlos E. Jimenez, Olga Russakovsky, Karthik Narasimhan
Abstract
We introduce CARETS, a systematic test suite to measure consistency and robustness of modern VQA models through a series of six fine-grained capability tests. In contrast to existing VQA test sets, CARETS features balanced question generation to create pairs of instances to test models, with each pair focusing on a specific capability such as rephrasing, logical symmetry or image obfuscation. We evaluate six modern VQA systems on CARETS and identify several actionable weaknesses in model comprehension, especially with concepts such as negation, disjunction, or hypernym invariance. Interestingly, even the most sophisticated models are sensitive to aspects such as swapping the order of terms in a conjunction or changing the number of answer choices mentioned in the question. We release CARETS to be used as an extensible tool for evaluating multi-modal model robustness. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Equivariant Similarity for Vision-Language Foundation ModelsTan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin et al.ICCV 2023 · 67 citations
- BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation PretrainingMinJun Kim, Seungwoo Song, Youhan Lee, Haneol Jang et al.AAAI 2024 · 10 citations
- Z-LaVI: Zero-Shot Language Solver Fueled by Visual ImaginationYue Yang, Wenlin Yao, Hongming Zhang, Xiaoyang Wang et al.EMNLP 2022 · 7 citations
- Understanding ME? Multimodal Evaluation for Fine-grained Visual CommonsenseZhecan Wang, Haoxuan You, Yicheng He, Wenhao Li et al.EMNLP 2022 · 2 citations
Builds on6
- A Study of Face Obfuscation in ImageNetKaiyu Yang, Jacqueline H. Yau, Li Fei-Fei, Jia Deng et al.ICML 2022 · 163 citations
- Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for TextNishtha Madaan, Inkit Padhi, Naveen Panwar, Diptikalyan SahaAAAI 2021 · 115 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
- Measuring CLEVRness: Black-box Testing of Visual Reasoning ModelsSpyridon Mouselinos, Henryk Michalewski, Mateusz MalinowskiICLR 2022 · 4 citations
- SQuINTing at VQA Models: Introspecting VQA Models With Sub-QuestionsRamprasaath R. Selvaraju, Purva Tendulkar, Devi Parikh, Eric Horvitz et al.CVPR 2020
Related papers
- Perception Matters: Detecting Perception Failures of VQA Models Using Metamorphic TestingYuanyuan Yuan, Shuai Wang, Mingyue Jiang, Tsong Yueh ChenCVPR 2021
- Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic EditingVedika Agarwal, Rakshith Shetty, Mario FritzCVPR 2020
- ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question AnsweringXiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang et al.CVPR 2026
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 198 citations
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan et al.NeurIPS 2024 · 27 citations
