Generative Bias for Robust Visual Question Answering
Jae-Won Cho, Dong-Jin Kim, Hyeonggon Ryu, In So Kweon
摘要
The task of Visual Question Answering (VQA) is known to be plagued by the issue of VQA models exploiting biases within the dataset to make its final prediction. Various previous ensemble based debiasing methods have been proposed where an additional model is purposefully trained to be biased in order to train a robust target model. However, these methods compute the bias for a model simply from the label statistics of the training data or from single modal branches. In this work, in order to better learn the bias a target VQA model suffers from, we propose a generative method to train the bias model directly from the target model, called GenB. In particular, GenB employs a generative network to learn the bias in the target model through a combination of the adversarial objective and knowledge distillation. We then debias our target model with GenB as a bias model, and show through extensive experiments the effects of our method on various VQA bias datasets including VQA-CP2, VQA-CP1, GQA-OOD, and VQA-CE, and show state-of-the-art results with the LXMERT architecture on VQA-CP2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question AnsweringJie Ma, Min Hu, Pinghui Wang, Wangchun Sun 等NeurIPS 2024 · 被引用 31 次
- VEGAS: Towards Visually Explainable and Grounded Artificial Social IntelligenceHao Li, Hao Fei, Zechao Hu, Zhengwei Yang 等AAAI 2025 · 被引用 6 次
- Towards Robust Visual Question Answering via Prompt-Driven Geometric HarmonizationYishu Liu, Jiawei Zhu, Congcong Wen, Guangming Lu 等AAAI 2025 · 被引用 3 次
- Language-Bias-Resilient Visual Question Answering via Adaptive Multi-Margin Collaborative DebiasingHuanjia Zhu, Shuyuan Zheng, Yishu Liu, Sudong Cai 等NeurIPS 2025 · 被引用 2 次
- Consistency of Compositional Generalization Across Multiple LevelsChuanhao Li, Zhen Li, Chenchen Jing, Xiaomeng Fan 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper14
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin 等ICCV 2019 · 被引用 288 次
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha 等NeurIPS 2020 · 被引用 163 次
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 被引用 136 次
- Overcoming Language Priors in VQA via Decomposed Linguistic RepresentationsChenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia 等AAAI 2020 · 被引用 115 次
相关 Paper
- Greedy Gradient Ensemble for Robust Visual Question AnsweringXinzhe Han, Shuhui Wang, Chi Su, Qingming Huang 等ICCV 2021 · 被引用 94 次
- Introspective Distillation for Robust Question AnsweringYulei Niu, Hanwang ZhangNeurIPS 2021 · 被引用 74 次
- Deconfounded Visual Question Generation with Causal InferenceJiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai 等ACM MM 2023 · 被引用 8 次
- X-GGM: Graph Generative Modeling for Out-of-distribution Generalization in Visual Question AnsweringJingjing Jiang, Ziyi Liu, Yifan Liu, Zhixiong Nan 等ACM MM 2021 · 被引用 17 次
- OAD-Promoter: Enhancing Zero-Shot VQA Using Large Language Models with Object Attribute DescriptionQuanxing Xu, Ling Zhou, Feifei Zhang, Rubing Huang 等AAAI 2026
