Perception Matters: Detecting Perception Failures of VQA Models Using Metamorphic Testing
Yuanyuan Yuan, Shuai Wang, Mingyue Jiang, Tsong Yueh Chen
摘要
Visual question answering (VQA) takes an image and a natural-language question as input and returns a naturallanguage answer. To date, VQA models are primarily assessed by their accuracy on high-level reasoning questions. Nevertheless, Given that perception tasks (e.g., recognizing objects) are the building blocks in the compositional process required by high-level reasoning, there is a demanding need to gain insights into how much of a problem lowlevel perception is. Inspired by the principles of software metamorphic testing, we introduce MetaVQA, a modelagnostic framework for benchmarking perception capability of VQA models. Given an image i, MetaVQA is able to synthesize a low-level perception question q. It then jointly transforms (i, q) to one or a set of sub-questions and subimages. MetaVQA checks whether the answer to (i, q) satisfies metamorphic relationships (MRs), denoting perception consistency, with the composed answers of transformed questions and images. Violating MRs denotes a failure of answering perception questions. MetaVQA successfully detects over 4.9 million perception failures made by popular VQA models with metamorphic testing. The state-of-the-art VQA models (e.g., the champion of VQA 2020 Challenge) suffer from perception consistency problems. In contrast, the Oscar VQA models, by using anchor points to align questions and images, show generally better consistency in perception tasks. We hope MetaVQA will revitalize interest in enhancing the low-level perceptual abilities of VQA models, a cornerstone of high-level reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- CCTEST: Testing and Repairing Code Completion SystemsZongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang 等ICSE 2023 · 被引用 49 次
- MDPFuzz: testing models solving Markov decision processesQi Pang, Yuanyuan Yuan, Shuai WangISSTA 2022 · 被引用 37 次
- Maintaining Reasoning Consistency in Compositional Visual Question AnsweringChenchen Jing, Yunde Jia, Yuwei Wu, Xinyu Liu 等CVPR 2022 · 被引用 27 次
- Pinolo: Detecting Logical Bugs in Database Management Systems with Approximate Query SynthesisZongyin Hao, Quanfeng Huang, Chengpeng Wang, Jianfeng Wang 等USENIX ATC 2023 · 被引用 26 次
- Unveiling Hidden DNN Defects with Decision-Based Metamorphic TestingYuanyuan Yuan, Qi Pang, Shuai WangASE 2022 · 被引用 16 次
它引用的顶会 Paper10
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 被引用 136 次
- Structure-invariant testing for machine translationPinjia He, Clara Meister, Zhendong SuICSE 2020 · 被引用 84 次
- Metamorphic Object Insertion for Testing Object Detection SystemsShuai Wang, Zhendong SuASE 2020 · 被引用 69 次
- MoVie: Revisiting Modulated Convolutions for Visual Counting and BeyondDuy-Kien Nguyen, Vedanuj Goswami, Xinlei ChenICLR 2021 · 被引用 12 次
相关 Paper
- Natural Test Generation for Precise Testing of Question Answering SoftwareQingchao Shen, Junjie Chen, Jie M. Zhang, Haoyu Wang 等ASE 2022 · 被引用 27 次
- Testing Your Question Answering Software via Asking RecursivelySongqiang Chen, Shuo Jin, Xiaoyuan XieASE 2021 · 被引用 37 次
- Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Li Niu, Liqing ZhangCVPR 2024
- Logical Implications for Visual Question Answering ConsistencySergio Tascon-Morales, Pablo Márquez-Neila, Raphael SznitmanCVPR 2023
- WebQA: Multihop and Multimodal QAYingshan Chang, Guihong Cao, Mridu Narang, Jianfeng Gao 等CVPR 2022 · 被引用 58 次
