Improving Commonsense in Vision-Language Models via Knowledge Graph Riddles
Shuquan Ye, Yujia Xie, Dongdong Chen, Yichong Xu, Lu Yuan, Chenguang Zhu, Jing Liao
摘要
This paper focuses on analyzing and improving the commonsense ability of recent popular vision-language (VL) models. Despite the great success, we observe that existing VL-models still lack commonsense knowledge/reasoning ability (e.g., "Lemons are sour"), which is a vital component towards artificial general intelligence. Through our analysis, we find one important reason is that existing large-scale VL datasets do not contain much commonsense knowledge, which motivates us to improve the commonsense of VL-models from the data perspective. Rather than collecting a new VL training dataset, we propose a more scalable strategy, i.e., "Data Augmentation with kNowledge graph linearization for CommonsensE capability" (DANCE). It can be viewed as one type of data augmentation technique, which can inject commonsense knowledge into existing VL datasets on the fly during training. More specifically, we leverage the commonsense knowledge graph (e.g., ConceptNet) and create variants of text description in VL datasets via bidirectional sub-graph sequentialization. For better commonsense evaluation, we further propose the first retrieval-based commonsense diagnostic benchmark. By conducting extensive experiments on some representative VL-models, we demonstrate that our DANCE technique is able to significantly improve the commonsense ability while maintaining the performance on vanilla retrieval tasks. The code and data are available at https:// github.com/ pleaseconnectwifi/ DANCE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Self-supervised Pre-training for Mirror DetectionJiaying Lin, Rynson W. H. LauICCV 2023 · 被引用 9 次
- Beyond Accuracy: Ensuring Correct Predictions With Correct RationalesTang Li, Mengmeng Ma, Xi PengNeurIPS 2024 · 被引用 6 次
- OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene UnderstandingYoujun Zhao, Jiaying Lin, Shuquan Ye, Qianshi Pang 等AAAI 2026 · 被引用 5 次
- Investigating Compositional Challenges in Vision-Language Models for Visual GroundingYunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie 等CVPR 2024 · 被引用 4 次
- Infer the Whole from a Glimpse of a Part: Keypoint-Based Knowledge Graph for Vehicle Re-IdentificationKai Lv, Yunlong Li, Zhuo Chen, Shuo Wang 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationJun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon 等EMNLP 2023
- CALM: Commen-Sense Knowledge Augmentation for Document Image UnderstandingQinyi Du, Qingqing Wang, Keqian Li, Jidong Tian 等ACM MM 2022 · 被引用 4 次
- Broaden the Vision: Geo-Diverse Visual Commonsense ReasoningDa Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng 等EMNLP 2021 · 被引用 32 次
- Retrieval Augmentation for Commonsense Reasoning: A Unified ApproachWenhao Yu, Chenguang Zhu, Zhihan Zhang, Shuohang Wang 等EMNLP 2022 · 被引用 10 次
