Visual Commonsense R-CNN
Tan Wang, Jianqiang Huang, Hanwang Zhang, Qianru Sun
Abstract
We present a novel unsupervised feature representation learning method, Visual Commonsense Region-based Convolutional Neural Network (VC R-CNN), to serve as an improved visual region encoder for high-level tasks such as captioning and VQA. Given a set of detected object regions in an image (e.g., using Faster R-CNN), like any other unsupervised feature learning methods (e.g., word2vec), the proxy training objective of VC R-CNN is to predict the contextual objects of a region. However, they are fundamentally different: the prediction of VC R-CNN is by using causal intervention: P (Y |do(X)), while others are by using the conventional likelihood: P (Y |X). This is also the core reason why VC R-CNN can learn "sense-making" knowledge like chair can be sat -while not just "common" co-occurrences such as chair is likely to exist if table is observed. We extensively apply VC R-CNN features in prevailing models of three popular tasks: Image Captioning, VQA, and VCR, and observe consistent performance boosts across them, achieving many new state-of-the-arts 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0aea1c81-688e-4cec-b002-0151e3e27accCited by top-tier papers88
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua et al.NeurIPS 2020 · 563 citations
- Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identificationYongming Rao, Guangyi Chen, Jiwen Lu, Jie ZhouICCV 2021 · 330 citations
- Interventional Few-Shot LearningZhongqi Yue, Hanwang Zhang, Qianru Sun, Xian-Sheng HuaNeurIPS 2020 · 284 citations
- Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender SystemTianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu et al.KDD 2021 · 246 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
Builds on6
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke et al.ICLR 2020 · 371 citations
- Learning to Collocate Neural Modules for Image CaptioningXu Yang, Hanwang Zhang, Jianfei CaiICCV 2019 · 84 citations
- Two Causal Principles for Improving Visual DialogJiaxin Qi, Yulei Niu, Jianqiang Huang, Hanwang ZhangCVPR 2020
Related papers
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression MatchingChi Chen, Peng Li, Maosong Sun, Yang LiuEMNLP 2022 · 7 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Deconfounded Visual Question Generation with Causal InferenceJiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai et al.ACM MM 2023 · 8 citations
- Hybrid Reasoning Network for Video-based Commonsense CaptioningWeijiang Yu, Jian Liang, Lei Ji, Lu Li et al.ACM MM 2021 · 8 citations
