Two Causal Principles for Improving Visual Dialog
Jiaxin Qi, Yulei Niu, Jianqiang Huang, Hanwang Zhang
Abstract
This paper unravels the design tricks adopted by us -the champion team MReaL-BDAI -for Visual Dialog Challenge 2019: two causal principles for improving Visual Dialog (VisDial). By "improving", we mean that they can promote almost every existing VisDial model to the stateof-the-art performance on the leader-board. Such a major improvement is only due to our careful inspection on the causality behind the model and data, finding that the community has overlooked two causalities in VisDial. Intuitively, Principle 1 suggests: we should remove the direct input of the dialog history to the answer model, otherwise a harmful shortcut bias will be introduced; Principle 2 says: there is an unobserved confounder for history, question, and answer, leading to spurious correlations from training data. In particular, to remove the confounder suggested in Principle 2, we propose several causal intervention algorithms, which make the training fundamentally different from the traditional likelihood estimation. Note that the two principles are model-agnostic, so they are applicable in any Vis-Dial model. The code is available at https://github . com/simpleshinobu/visdial-principles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers47
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua et al.NeurIPS 2020 · 563 citations
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 533 citations
- Interventional Few-Shot LearningZhongqi Yue, Hanwang Zhang, Qianru Sun, Xian-Sheng HuaNeurIPS 2020 · 284 citations
- Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender SystemTianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu et al.KDD 2021 · 246 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
Builds on3
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke et al.ICLR 2020 · 371 citations
- Visual Commonsense R-CNNTan Wang, Jianqiang Huang, Hanwang Zhang, Qianru SunCVPR 2020
- Unbiased Scene Graph Generation From Biased TrainingKaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi et al.CVPR 2020
Related papers
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- Unsupervised and Pseudo-Supervised Vision-Language Alignment in Visual DialogFeilong Chen, Duzhen Zhang, Xiuyi Chen, Jing Shi et al.ACM MM 2022 · 10 citations
- History for Visual Dialog: Do we really need it?Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas et al.ACL 2020 · 8 citations
- Making History Matter: History-Advantage Sequence Training for Visual DialogTianhao Yang, Zheng-Jun Zha, Hanwang ZhangICCV 2019 · 71 citations
- V^2Dial: Unification of Video and Visual Dialog via Multimodal ExpertsAdnen Abdessaied, Anna Rohrbach, Marcus Rohrbach, Andreas BullingCVPR 2025
