Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-Modal Knowledge Transfer
Woojeong Jin, Dong-Ho Lee, Chenguang Zhu, Jay Pujara, Xiang Ren
Abstract
Pre-trained language models are still far from human performance in tasks that need understanding of properties (e.g. appearance, measurable quantity) and affordances of everyday objects in the real world since the text lacks such information due to reporting bias. In this work, we study whether integrating visual knowledge into a language model can fill the gap. We investigate two types of knowledge transfer: (1) text knowledge transfer using image captions that may contain enriched visual knowledge and (2) cross-modal knowledge transfer using both images and captions with vision-language training objectives. On 5 downstream tasks that may need visual knowledge to solve the problem, we perform extensive empirical comparisons over the presented objectives. Our experiments show that visual knowledge transfer can improve performance in both low-resource and fully supervised settings. 1 * Authors contributed equally. 1 https://github.com/INK-USC/CMKT Interesting facts about orange ! 1. Orange elevates mood levels. 2. Orange are often grown in the Mediterranean. 3. Oranges facing the sunnier tend to be sweeter. Human Typical facts about orange … 1. Orange is a shape of circle. 2. Orange is a color of orange.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93539418-be6b-4ed1-9774-e43aab4f78eaCited by top-tier papers4
- Learning to Imagine: Visually-Augmented Natural Language GenerationTianyi Tang, Yushuo Chen, Yifan Du, Junyi Li et al.ACL 2023 · 7 citations
- MPCHAT: Towards Multimodal Persona-Grounded ConversationJaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee KimACL 2023 · 4 citations
- ReSeeding Latent States for Sequential Language UnderstandingStéphane Aroca-Ouellette, Katharina von der Wense, Alessandro RonconeEMNLP 2025 · 1 citation
- Transparent and Coherent Procedural Mistake DetectionShane Storks, Itamar Bar-Yossef, Yayuan Li, Zheyuan Zhang et al.EMNLP 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
Related papers
- Visually-Augmented Language ModelingWeizhi Wang, Li Dong, Hao Cheng, Haoyu Song et al.ICLR 2023 · 5 citations
- Z-LaVI: Zero-Shot Language Solver Fueled by Visual ImaginationYue Yang, Wenlin Yao, Hongming Zhang, Xiaoyang Wang et al.EMNLP 2022 · 7 citations
- Transitional Adaptation of Pretrained Models for Visual StorytellingYoungjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim et al.CVPR 2021
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin et al.CVPR 2021
