FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
Chengyue Huang, Brisa Maneechotesuwan, Shivang Chopra, Zsolt Kira
摘要
Visual question answering (VQA) systems face significant challenges when adapting to real-world data shifts, especially in multi-modal contexts. While robust fine-tuning strategies are essential for maintaining performance across in-distribution (ID) and out-of-distribution (OOD) scenarios, current evaluation settings are primarily unimodal or particular to some types of OOD, offering limited insight into the complexities of multi-modal contexts. In this work, we propose a new benchmark FRAMES-VQA (Fine-Tuning Robustness Across Multi-Modal Shifts in VQA) for evaluating robust fine-tuning for VQA tasks. We utilize ten existing VQA benchmarks, including VQAv2, IV-VQA, VQA-CP, OK-VQA and others, and categorize them into ID, near and far OOD datasets covering uni-modal, multi-modal and adversarial distribution shifts. We first conduct a comprehensive comparison of existing robust fine-tuning methods. We then quantify the distribution shifts by calculating the Mahalanobis distance using uni-modal and multimodal embeddings extracted from various models. Further, we perform an extensive analysis to explore the interactions between uni-and multi-modal shifts as well as modality importance for ID and OOD samples. These analyses offer valuable guidance on developing more robust fine-tuning methods to handle multi-modal distribution shifts. The code is available at https://github.com/ chengyuehuang511/FRAMES-VQA. * Equal contribution. and clipart, challenging models to generalize across different styles and representations. Similarly, various ImageNet variants [10, 22, 39, 50] introduce shifts through image variations, adversarial examples, rendering transformations, and changes in texture or background. Collectively, these datasets provide a comprehensive framework for assessing how well models withstand visual distribution changes. While robust fine-tuning algorithms are widely examined under distribution shifts in a single modality (images), few studies have explored robust fine-tuning for VQA tasks, where distribution shifts are multi-modal and models must adapt to variations across both visual and textual inputs. Apart from visual shift [1], there are question shifts [15, 41] involving variations in phrasing, structure, or vocabulary, as well as answer shifts [2] with changes in answer distributions such as frequency and formatting. Beyond uni-modal shift, these variations may occur simultaneously across visual, question, and answer inputs [9, 31, 42, 44, 49] , posing an even greater challenge as models must generalize across complex, combined shifts. Therefore, we build upon our preliminary exploration [25] and propose a benchmark FRAMES-VQA (Fine-Tuning Robustness Across Multi-Modal Shifts in VQA) to systematically evaluate the robustness of fine-tuning in VQA task. We leverage ten existing VQA datasets and categorize distribution shifts into uni-modal and multi-modal types, quantified by Mahalanobis distance across various backbones to capture both near and far OOD scenarios. We conduct a comprehensive comparison of the existing robust fine-tuning baselines on ID and OOD performance using the benchmark. Furthermore, we analyze shift scores and modality importance across fine-tuning methods. To summarize, our contributions are: • We propose FRAMES-VQA for evaluating robust finetuning in VQA, including ten VQA datasets categorized by uni-modal (e.g., image, question) and multi-modal shifts. We quantify dataset shifts under different modalities using Mahalanobis distance and embeddings from different backbones. • We perform an in-depth comparison of robust fine-tuning This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action GeneralizationChengyue Huang, Mellon M. Zhang, Robert Azarcon, Glen Chou 等CVPR 2026 · 被引用 8 次
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question AnsweringJian Lan, Zhicheng Liu, Udo Schlegel, Raoyuan Zhao 等ICLR 2026 · 被引用 2 次
- EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic RetrievalJiashi Lin, Changhong Jiang, Xiangru Lin, Ruifei Zhang 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
相关 Paper
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma 等EMNLP 2021 · 被引用 18 次
- Directional Gradient Projection for Robust Fine-Tuning of Foundation ModelsChengyue Huang, Junjiao Tian, Brisa Maneechotesuwan, Shivang Chopra 等ICLR 2025
- Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual LearningZijian Gao, Zicheng Sun, Xingxing Zhang, Kele Xu 等CVPR 2026
- Do-GOOD: Towards Distribution Shift Evaluation for Pre-Trained Visual Document Understanding ModelsJiabang He, Yi Hu, Lei Wang, Xing Xu 等SIGIR 2023 · 被引用 4 次
- AQuA: Toward Strategic Response Generation for Ambiguous Visual QuestionsJihyoung Jang, Hyounghun KimICLR 2026 · 被引用 1 次
