Improving Commonsense in Vision-Language Models via Knowledge Graph Riddles
Shuquan Ye, Yujia Xie, Dongdong Chen, Yichong Xu, Lu Yuan, Chenguang Zhu, Jing Liao
Abstract
This paper focuses on analyzing and improving the commonsense ability of recent popular vision-language (VL) models. Despite the great success, we observe that existing VL-models still lack commonsense knowledge/reasoning ability (e.g., "Lemons are sour"), which is a vital component towards artificial general intelligence. Through our analysis, we find one important reason is that existing large-scale VL datasets do not contain much commonsense knowledge, which motivates us to improve the commonsense of VL-models from the data perspective. Rather than collecting a new VL training dataset, we propose a more scalable strategy, i.e., "Data Augmentation with kNowledge graph linearization for CommonsensE capability" (DANCE). It can be viewed as one type of data augmentation technique, which can inject commonsense knowledge into existing VL datasets on the fly during training. More specifically, we leverage the commonsense knowledge graph (e.g., ConceptNet) and create variants of text description in VL datasets via bidirectional sub-graph sequentialization. For better commonsense evaluation, we further propose the first retrieval-based commonsense diagnostic benchmark. By conducting extensive experiments on some representative VL-models, we demonstrate that our DANCE technique is able to significantly improve the commonsense ability while maintaining the performance on vanilla retrieval tasks. The code and data are available at https:// github.com/ pleaseconnectwifi/ DANCE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc68ef14-2f65-4be5-b721-9d6151bf10d3Cited by top-tier papers7
- Self-supervised Pre-training for Mirror DetectionJiaying Lin, Rynson W. H. LauICCV 2023 · 9 citations
- Beyond Accuracy: Ensuring Correct Predictions With Correct RationalesTang Li, Mengmeng Ma, Xi PengNeurIPS 2024 · 6 citations
- OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene UnderstandingYoujun Zhao, Jiaying Lin, Shuquan Ye, Qianshi Pang et al.AAAI 2026 · 5 citations
- Investigating Compositional Challenges in Vision-Language Models for Visual GroundingYunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie et al.CVPR 2024 · 4 citations
- Infer the Whole from a Glimpse of a Part: Keypoint-Based Knowledge Graph for Vehicle Re-IdentificationKai Lv, Yunlong Li, Zhuo Chen, Shuo Wang et al.AAAI 2025 · 1 citation
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
- DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationJun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon et al.EMNLP 2023
- CALM: Commen-Sense Knowledge Augmentation for Document Image UnderstandingQinyi Du, Qingqing Wang, Keqian Li, Jidong Tian et al.ACM MM 2022 · 4 citations
- Broaden the Vision: Geo-Diverse Visual Commonsense ReasoningDa Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng et al.EMNLP 2021 · 32 citations
- Retrieval Augmentation for Commonsense Reasoning: A Unified ApproachWenhao Yu, Chenguang Zhu, Zhihan Zhang, Shuohang Wang et al.EMNLP 2022 · 10 citations
