Contrast and Classify: Training Robust VQA Models
Yash Kant, Abhinav Moudgil, Dhruv Batra, Devi Parikh, Harsh Agrawal
Abstract
Recent Visual Question Answering (VQA) models have shown impressive performance on the VQA benchmark but remain sensitive to small linguistic variations in input questions. Existing approaches address this by augmenting the dataset with question paraphrases from visual question generation models or adversarial perturbations. These approaches use the combined data to learn an answer classifier by minimizing the standard cross-entropy loss. To more effectively leverage augmented data, we build on the recent success in contrastive learning. We propose a novel training paradigm (ConClaT) that optimizes both cross-entropy and contrastive losses. The contrastive loss encourages representations to be robust to linguistic variations in questions while the cross-entropy loss preserves the discriminative power of representations for answer prediction.We find that optimizing both losses – either alternately or jointly – is key to effective training. On the VQA-Rephrasings [44] benchmark, which measures the VQA model’s answer consistency across human paraphrases of a question, ConClaT improves Consensus Score by 1.63% over an improved baseline. In addition, on the standard VQA 2.0 benchmark, we improve the VQA accuracy by 0.78% overall. We also show that ConClaT is agnostic to the type of data-augmentation strategy used.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 873e627c-86d4-4715-b970-a7ede9d2b46aCited by top-tier papers3
- FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual CompressionBo Tong, Bokai Lai, Yiyi Zhou, Gen Luo et al.CVPR 2025
- Multi-Modal Representation Learning with Text-Driven Soft MasksJaeyoo Park, Bohyung HanCVPR 2023
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMsAbhirama Subramanyam Penamakuri, Navlika Singh, Piyush Arora, Anand MishraEMNLP 2025
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
Related papers
- Multimodal Hypothetical Summary for Retrieval-based Multi-image Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin et al.AAAI 2025 · 1 citation
- Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMsParsa Hejabi, Elnaz Rahmati, Alireza Salkhordeh Ziabari, Morteza DehghaniACL 2026 · 1 citation
- Robust Cross-Modal Representation Learning with Progressive Self-DistillationAlex Andonian, Shixing Chen, Raffay HamidCVPR 2022 · 43 citations
- Balanced Adversarial Training: Balancing Tradeoffs between Fickleness and Obstinacy in NLP ModelsHannah Chen, Yangfeng Ji, David E. EvansEMNLP 2022 · 4 citations
- Logical Implications for Visual Question Answering ConsistencySergio Tascon-Morales, Pablo Márquez-Neila, Raphael SznitmanCVPR 2023
