FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
Chengyue Huang, Brisa Maneechotesuwan, Shivang Chopra, Zsolt Kira
Abstract
Visual question answering (VQA) systems face significant challenges when adapting to real-world data shifts, especially in multi-modal contexts. While robust fine-tuning strategies are essential for maintaining performance across in-distribution (ID) and out-of-distribution (OOD) scenarios, current evaluation settings are primarily unimodal or particular to some types of OOD, offering limited insight into the complexities of multi-modal contexts. In this work, we propose a new benchmark FRAMES-VQA (Fine-Tuning Robustness Across Multi-Modal Shifts in VQA) for evaluating robust fine-tuning for VQA tasks. We utilize ten existing VQA benchmarks, including VQAv2, IV-VQA, VQA-CP, OK-VQA and others, and categorize them into ID, near and far OOD datasets covering uni-modal, multi-modal and adversarial distribution shifts. We first conduct a comprehensive comparison of existing robust fine-tuning methods. We then quantify the distribution shifts by calculating the Mahalanobis distance using uni-modal and multimodal embeddings extracted from various models. Further, we perform an extensive analysis to explore the interactions between uni-and multi-modal shifts as well as modality importance for ID and OOD samples. These analyses offer valuable guidance on developing more robust fine-tuning methods to handle multi-modal distribution shifts. The code is available at https://github.com/ chengyuehuang511/FRAMES-VQA. * Equal contribution. and clipart, challenging models to generalize across different styles and representations. Similarly, various ImageNet variants [10, 22, 39, 50] introduce shifts through image variations, adversarial examples, rendering transformations, and changes in texture or background. Collectively, these datasets provide a comprehensive framework for assessing how well models withstand visual distribution changes. While robust fine-tuning algorithms are widely examined under distribution shifts in a single modality (images), few studies have explored robust fine-tuning for VQA tasks, where distribution shifts are multi-modal and models must adapt to variations across both visual and textual inputs. Apart from visual shift [1], there are question shifts [15, 41] involving variations in phrasing, structure, or vocabulary, as well as answer shifts [2] with changes in answer distributions such as frequency and formatting. Beyond uni-modal shift, these variations may occur simultaneously across visual, question, and answer inputs [9, 31, 42, 44, 49] , posing an even greater challenge as models must generalize across complex, combined shifts. Therefore, we build upon our preliminary exploration [25] and propose a benchmark FRAMES-VQA (Fine-Tuning Robustness Across Multi-Modal Shifts in VQA) to systematically evaluate the robustness of fine-tuning in VQA task. We leverage ten existing VQA datasets and categorize distribution shifts into uni-modal and multi-modal types, quantified by Mahalanobis distance across various backbones to capture both near and far OOD scenarios. We conduct a comprehensive comparison of the existing robust fine-tuning baselines on ID and OOD performance using the benchmark. Furthermore, we analyze shift scores and modality importance across fine-tuning methods. To summarize, our contributions are: • We propose FRAMES-VQA for evaluating robust finetuning in VQA, including ten VQA datasets categorized by uni-modal (e.g., image, question) and multi-modal shifts. We quantify dataset shifts under different modalities using Mahalanobis distance and embeddings from different backbones. • We perform an in-depth comparison of robust fine-tuning This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action GeneralizationChengyue Huang, Mellon M. Zhang, Robert Azarcon, Glen Chou et al.CVPR 2026 · 8 citations
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question AnsweringJian Lan, Zhicheng Liu, Udo Schlegel, Raoyuan Zhao et al.ICLR 2026 · 2 citations
- EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic RetrievalJiashi Lin, Changhong Jiang, Xiangru Lin, Ruifei Zhang et al.CVPR 2026 · 2 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma et al.EMNLP 2021 · 18 citations
- Directional Gradient Projection for Robust Fine-Tuning of Foundation ModelsChengyue Huang, Junjiao Tian, Brisa Maneechotesuwan, Shivang Chopra et al.ICLR 2025
- Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual LearningZijian Gao, Zicheng Sun, Xingxing Zhang, Kele Xu et al.CVPR 2026
- Do-GOOD: Towards Distribution Shift Evaluation for Pre-Trained Visual Document Understanding ModelsJiabang He, Yi Hu, Lei Wang, Xing Xu et al.SIGIR 2023 · 4 citations
- AQuA: Toward Strategic Response Generation for Ambiguous Visual QuestionsJihyoung Jang, Hyounghun KimICLR 2026 · 1 citation
