MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question Answering
Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang
Abstract
While progress has been made on the visual question answering leaderboards, models often utilize spurious correlations and priors in datasets under the i.i.d. setting. As such, evaluation on out-of-distribution (OOD) test samples has emerged as a proxy for generalization. In this paper, we present MUTANT, a training paradigm that exposes the model to perceptually similar, yet semantically distinct mutations of the input, to improve OOD generalization, such as the VQA-CP challenge. Under this paradigm, models utilize a consistency-constrained training objective to understand the effect of semantic changes in input (question-image pair) on the output (answer). Unlike existing methods on VQA-CP, MUTANT does not rely on the knowledge about the nature of train and test answer distributions. MUTANT establishes a new state-ofthe-art accuracy on VQA-CP with a 10.57% improvement. Our work opens up avenues for the use of semantic input mutations for OOD generalization in question answering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 083a24ed-5587-4ae9-b0a1-9627db209f3bCited by top-tier papers27
- VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic PhenomenaLetitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank et al.ACL 2022 · 147 citations
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu et al.NeurIPS 2021 · 102 citations
- Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question AnsweringCorentin Dancette, Rémi Cadène, Damien Teney, Matthieu CordICCV 2021 · 95 citations
- Introspective Distillation for Robust Question AnsweringYulei Niu, Hanwang ZhangNeurIPS 2021 · 74 citations
- Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question AnsweringMan Luo, Yankai Zeng, Pratyay Banerjee, Chitta BaralEMNLP 2021 · 53 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Self-Supervised Knowledge Triplet Learning for Zero-Shot Question AnsweringPratyay Banerjee, Chitta BaralEMNLP 2020 · 20 citations
Related papers
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha et al.NeurIPS 2020 · 163 citations
- X-GGM: Graph Generative Modeling for Out-of-distribution Generalization in Visual Question AnsweringJingjing Jiang, Ziyi Liu, Yifan Liu, Zhixiong Nan et al.ACM MM 2021 · 17 citations
- Unshuffling Data for Improved Generalization in Visual Question AnsweringDamien Teney, Ehsan Abbasnejad, Anton van den HengelICCV 2021 · 84 citations
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma et al.EMNLP 2021 · 18 citations
- Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic EditingVedika Agarwal, Rakshith Shetty, Mario FritzCVPR 2020
