Introspective Distillation for Robust Question Answering
Yulei Niu, Hanwang Zhang
摘要
Question answering (QA) models are well-known to exploit data bias, e.g., the language prior in visual QA and the position bias in reading comprehension. Recent debiasing methods achieve good out-of-distribution (OOD) generalizability with a considerable sacrifice of the in-distribution (ID) performance. Therefore, they are only applicable in domains where the test distribution is known in advance. In this paper, we present a novel debiasing method called Introspective Distillation (IntroD) to make the best of both worlds for QA. Our key technical contribution is to blend the inductive bias of OOD and ID by introspecting whether a training sample fits in the factual ID world or the counterfactual OOD one. Experiments on visual QA datasets VQA v2, VQA-CP, and reading comprehension dataset SQuAD demonstrate that our proposed IntroD maintains the competitive OOD performance compared to other debiasing methods, while sacrificing little or even achieving better ID performance compared to the non-debiasing ones. Question answering (QA), which requires machines to answer questions given a context, is one of the most fundamental AI tasks. Popular contexts are vision (e.g., image for VQA [10] ) and natural language (e.g., passage for extractive QA [38] ). A common observation is that QA models prefer to over-exploit the training bias, which bypasses the context comprehension for a shortcut answer. For example, by only using the linguistic correlations between questions and answers, VQA models can answer most questions correctly [22, 7, 10, 27] . Similarly, extractive QA models may use the spurious positional cues to locate the answer in the passage [30] . As a result, QA models that have already achieved strong in-distribution (ID) performance may inevitably fail in out-of-distribution (OOD) test scenarios, regardless of the scale of training data and models [20, 30, 50] . Recently, several debiasing methods aim to close the gap between the ID and OOD performances [12, 17, 13, 35] . However, many of them hold the assumption that the training and test distributions are very different or even reversed, e.g., if there are more "yes" answers in training, there must be more "no" answers in testing. As a result, these methods encounter a severe performance drop under the ID evaluation, although they significantly outperform non-debiasing baselines in terms of OOD performance. An interesting observation from Figure 1 is that non-debiasing methods (circles) obtain high ID but low OOD performance, while debiasing methods (squares) achieve high OOD but low ID performance. This observation motivates us to ask: can we make the best of both worlds? 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Respecting Transfer Gap in Knowledge DistillationYulei Niu, Long Chen, Chang Zhou, Hanwang ZhangNeurIPS 2022 · 被引用 36 次
- Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite GraphsKaifeng Gao, Long Chen, Yulei Niu, Jian Shao 等CVPR 2022 · 被引用 34 次
- Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question AnsweringJie Ma, Min Hu, Pinghui Wang, Wangchun Sun 等NeurIPS 2024 · 被引用 31 次
- Interpolative Distillation for Unifying Biased and Debiased RecommendationSihao Ding, Fuli Feng, Xiangnan He, Jinqiu Jin 等SIGIR 2022 · 被引用 27 次
- MAP: Towards Balanced Generalization of IID and OOD through Model-Agnostic AdaptersMin Zhang, Junkun Yuan, Yue He, Wenbin Li 等ICCV 2023 · 被引用 21 次
它引用的顶会 Paper16
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin 等ICCV 2019 · 被引用 288 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
- Self-Knowledge Distillation with Progressive Refinement of TargetsKyungyul Kim, Byeongmoon Ji, Doyoung Yoon, Sangheum HwangICCV 2021 · 被引用 251 次
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha 等NeurIPS 2020 · 被引用 163 次
相关 Paper
- Generative Bias for Robust Visual Question AnsweringJae-Won Cho, Dong-Jin Kim, Hyeonggon Ryu, In So KweonCVPR 2023
- Counterfactual VQA: A Cause-Effect Look at Language BiasYulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu 等CVPR 2021
- Greedy Gradient Ensemble for Robust Visual Question AnsweringXinzhe Han, Shuhui Wang, Chi Su, Qingming Huang 等ICCV 2021 · 被引用 94 次
- DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsShengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang 等ACM MM 2020 · 被引用 66 次
- Roses Are Red, Violets Are Blue... but Should VQA Expect Them To?Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian WolfCVPR 2021
