Domain-Robust VQA With Diverse Datasets and Methods but No Target Labels
Mingda Zhang, Tristan Maidment, Ahmad Diab, Adriana Kovashka, Rebecca Hwa
摘要
The observation that computer vision methods overfit to dataset specifics has inspired diverse attempts to make object recognition models robust to domain shifts. However, similar work on domain-robust visual question answering methods is very limited. Domain adaptation for VQA differs from adaptation for object recognition due to additional complexity: VQA models handle multimodal inputs, methods contain multiple steps with diverse modules resulting in complex optimization, and answer spaces in different datasets are vastly different. To tackle these challenges, we first quantify domain shifts between popular VQA datasets, in both visual and textual space. To disentangle shifts between datasets arising from different modalities, we also construct synthetic shifts in the image and question domains separately. Second, we test the robustness of different families of VQA methods (classic two-stream, transformer, and neuro-symbolic methods) to these shifts. Third, we test the applicability of existing domain adaptation methods and devise a new one to bridge VQA domain gaps, adjusted to specific VQA models. To emulate the setting of real-world generalization, we focus on unsupervised domain adaptation and the open-ended classification task formulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Sim VQA: Exploring Simulated Environments for Visual Question AnsweringPaola Cascante-Bonilla, Hui Wu, Letao Wang, Rogério Feris 等CVPR 2022 · 被引用 28 次
- Multi-Domain Lifelong Visual Question Answering via Self-Critical DistillationMingrui Lao, Nan Pu, Yu Liu, Zhun Zhong 等ACM MM 2023 · 被引用 5 次
- Crossing the Gap: Domain Generalization for Image CaptioningYuchen Ren, Zhendong Mao, Shancheng Fang, Yan Lu 等CVPR 2023
- FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question AnsweringChengyue Huang, Brisa Maneechotesuwan, Shivang Chopra, Zsolt KiraCVPR 2025
- Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual ReasoningZhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski 等CVPR 2023
它引用的顶会 Paper13
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- Larger Norm More Transferable: An Adaptive Feature Norm Approach for Unsupervised Domain AdaptationRuijia Xu, Guanbin Li, Jihan Yang, Liang LinICCV 2019 · 被引用 563 次
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu 等NeurIPS 2020 · 被引用 561 次
- Episodic Training for Domain GeneralizationDa Li, Jianshu Zhang, Yongxin Yang, Cong Liu 等ICCV 2019 · 被引用 488 次
- Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization Without Accessing Target Domain DataXiangyu Yue, Yang Zhang, Sicheng Zhao, Alberto L. Sangiovanni-Vincentelli 等ICCV 2019 · 被引用 462 次
相关 Paper
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma 等EMNLP 2021 · 被引用 18 次
- When Open-Vocabulary Visual Question Answering Meets Causal Adapter: Benchmark and ApproachFeifei Zhang, Zhaoyi Zhang, Xi Zhang, Changsheng XuAAAI 2025
- Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic EditingVedika Agarwal, Rakshith Shetty, Mario FritzCVPR 2020
- KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQAKenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta 等CVPR 2021
- Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQAWentao Mo, Yang LiuAAAI 2024 · 被引用 30 次
