Introspective Distillation for Robust Question Answering
Yulei Niu, Hanwang Zhang
Abstract
Question answering (QA) models are well-known to exploit data bias, e.g., the language prior in visual QA and the position bias in reading comprehension. Recent debiasing methods achieve good out-of-distribution (OOD) generalizability with a considerable sacrifice of the in-distribution (ID) performance. Therefore, they are only applicable in domains where the test distribution is known in advance. In this paper, we present a novel debiasing method called Introspective Distillation (IntroD) to make the best of both worlds for QA. Our key technical contribution is to blend the inductive bias of OOD and ID by introspecting whether a training sample fits in the factual ID world or the counterfactual OOD one. Experiments on visual QA datasets VQA v2, VQA-CP, and reading comprehension dataset SQuAD demonstrate that our proposed IntroD maintains the competitive OOD performance compared to other debiasing methods, while sacrificing little or even achieving better ID performance compared to the non-debiasing ones. Question answering (QA), which requires machines to answer questions given a context, is one of the most fundamental AI tasks. Popular contexts are vision (e.g., image for VQA [10] ) and natural language (e.g., passage for extractive QA [38] ). A common observation is that QA models prefer to over-exploit the training bias, which bypasses the context comprehension for a shortcut answer. For example, by only using the linguistic correlations between questions and answers, VQA models can answer most questions correctly [22, 7, 10, 27] . Similarly, extractive QA models may use the spurious positional cues to locate the answer in the passage [30] . As a result, QA models that have already achieved strong in-distribution (ID) performance may inevitably fail in out-of-distribution (OOD) test scenarios, regardless of the scale of training data and models [20, 30, 50] . Recently, several debiasing methods aim to close the gap between the ID and OOD performances [12, 17, 13, 35] . However, many of them hold the assumption that the training and test distributions are very different or even reversed, e.g., if there are more "yes" answers in training, there must be more "no" answers in testing. As a result, these methods encounter a severe performance drop under the ID evaluation, although they significantly outperform non-debiasing baselines in terms of OOD performance. An interesting observation from Figure 1 is that non-debiasing methods (circles) obtain high ID but low OOD performance, while debiasing methods (squares) achieve high OOD but low ID performance. This observation motivates us to ask: can we make the best of both worlds? 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Respecting Transfer Gap in Knowledge DistillationYulei Niu, Long Chen, Chang Zhou, Hanwang ZhangNeurIPS 2022 · 36 citations
- Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite GraphsKaifeng Gao, Long Chen, Yulei Niu, Jian Shao et al.CVPR 2022 · 34 citations
- Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question AnsweringJie Ma, Min Hu, Pinghui Wang, Wangchun Sun et al.NeurIPS 2024 · 31 citations
- Interpolative Distillation for Unifying Biased and Debiased RecommendationSihao Ding, Fuli Feng, Xiangnan He, Jinqiu Jin et al.SIGIR 2022 · 27 citations
- MAP: Towards Balanced Generalization of IID and OOD through Model-Agnostic AdaptersMin Zhang, Junkun Yuan, Yue He, Wenbin Li et al.ICCV 2023 · 21 citations
Builds on16
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin et al.ICCV 2019 · 288 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
- Self-Knowledge Distillation with Progressive Refinement of TargetsKyungyul Kim, Byeongmoon Ji, Doyoung Yoon, Sangheum HwangICCV 2021 · 251 citations
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha et al.NeurIPS 2020 · 163 citations
Related papers
- Generative Bias for Robust Visual Question AnsweringJae-Won Cho, Dong-Jin Kim, Hyeonggon Ryu, In So KweonCVPR 2023
- Counterfactual VQA: A Cause-Effect Look at Language BiasYulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu et al.CVPR 2021
- Greedy Gradient Ensemble for Robust Visual Question AnsweringXinzhe Han, Shuhui Wang, Chi Su, Qingming Huang et al.ICCV 2021 · 94 citations
- DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsShengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang et al.ACM MM 2020 · 66 citations
- Roses Are Red, Violets Are Blue... but Should VQA Expect Them To?Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian WolfCVPR 2021
