DeVLBert: Learning Deconfounded Visio-Linguistic Representations
Shengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang, Zhou Zhao, Jianke Zhu, Jin Yu, Hongxia Yang, Fei Wu
Abstract
In this paper, we propose to investigate the problem of out-of-domain visio-linguistic pretraining, where the pretraining data distribution differs from that of downstream data on which the pretrained model will be fine-tuned. Existing methods for this problem are purely likelihood-based, leading to the spurious correlations and hurt the generalization ability when transferred to out-of-domain downstream tasks. By spurious correlation, we mean that the conditional probability of one token (object or word) given another one can be high (due to the dataset biases) without robust (causal) relationships between them. To mitigate such dataset biases, we propose a Deconfounded Visio-Linguistic Bert framework, abbreviated as DeVLBert, to perform intervention-based learning. We borrow the idea of the backdoor adjustment from the research field of causality and propose several neural-network based architectures for Bert-style out-of-domain pretraining. The quantitative results on three downstream tasks, Image Retrieval (IR), Zero-shot IR, and Visual Question Answering, show the effectiveness of DeVLBert by boosting generalization ability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe63c283-07e4-4bd8-ac81-af29ae4a97a5Cited by top-tier papers36
- Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender SystemTianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu et al.KDD 2021 · 246 citations
- CauseRec: Counterfactual User Sequence Synthesis for Sequential RecommendationShengyu Zhang, Dong Yao, Zhou Zhao, Tat-Seng Chua et al.SIGIR 2021 · 118 citations
- Re4: Learning to Re-contrast, Re-attend, Re-construct for Multi-interest RecommendationShengyu Zhang, Lingxiao Yang, Dong Yao, Yujie Lu et al.WWW 2022 · 67 citations
- Should Graph Convolution Trust Neighbors? A Simple Causal Inference MethodFuli Feng, Weiran Huang, Xiangnan He, Xin Xin et al.SIGIR 2021 · 66 citations
- DocTr: Document Image Transformer for Geometric Unwarping and Illumination CorrectionHao Feng, Yuechen Wang, Wengang Zhou, Jiajun Deng et al.ACM MM 2021 · 66 citations
Builds on8
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Causally Denoise Word Embeddings Using Half-Sibling RegressionZekun Yang, Tianlin LiuAAAI 2020 · 7 citations
- Visual Commonsense R-CNNTan Wang, Jianqiang Huang, Hanwang Zhang, Qianru SunCVPR 2020
Related papers
- Show, Deconfound and Tell: Image Captioning with Causal InferenceBing Liu, Dong Wang, Xu Yang, Yong Zhou et al.CVPR 2022 · 66 citations
- Causal Attention for Vision-Language TasksXu Yang, Hanwang Zhang, Guojun Qi, Jianfei CaiCVPR 2021
- Debiasing NLU Models via Causal Intervention and Counterfactual ReasoningBing Tian, Yixin Cao, Yong Zhang, Chunxiao XingAAAI 2022 · 45 citations
- Debiased Fine-Tuning for Vision-Language Models by Prompt RegularizationBeier Zhu, Yulei Niu, Saeil Lee, Minhoe Hur et al.AAAI 2023 · 34 citations
- CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual EditingHaoxiang Cao, Chaoqun Wang, Yongwen Lai, Shaobo Min et al.ACM MM 2025 · 1 citation
