Explainable Semantic Space by Grounding Language to Vision with Cross-Modal Contrastive Learning
Yizhen Zhang, Minkyu Choi, Kuan Han, Zhongming Liu
Abstract
In natural language processing, most models try to learn semantic representations merely from texts. The learned representations encode the "distributional semantics" but fail to connect to any knowledge about the physical world. In contrast, humans learn language by grounding concepts in perception and action and the brain encodes "grounded semantics" for cognition. Inspired by this notion and recent work in vision-language learning, we design a two-stream model for grounding language learning in vision. The model includes a VGG-based visual stream and a Bert-based language stream. The two streams merge into a joint representational space. Through cross-modal contrastive learning, the model first learns to align visual and language representations with the MS COCO dataset. The model further learns to retrieve visual objects with language queries through a cross-modal attention module and to infer the visual relations between the retrieved objects through a bilinear operator with the Visual Genome dataset. After training, the model's language stream is a stand-alone language model capable of embedding concepts in a visually grounded semantic space. This semantic space manifests principal dimensions explainable with human intuition and neurobiological knowledge. Word embeddings in this semantic space are predictive of human-defined norms of semantic features and are segregated into perceptually distinctive clusters. Furthermore, the visually grounded language model also enables compositional language understanding based on visual knowledge and multimodal image search with queries based on images, texts, or their combinations. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e94150c4-6cf6-4679-9f9e-dfe0a586b260Cited by top-tier papers4
- ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single ModelJialong Zuo, Yongtai Deng, Mengdan Tan, Rui Jin et al.NeurIPS 2025 · 11 citations
- Mind Artist: Creating Artistic Snapshots with Human ThoughtJiaxuan Chen, Yu Qi, Yueming Wang, Gang PanCVPR 2024
- Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language ModelsDavide Berasi, Matteo Farina, Massimiliano Mancini, Elisa Ricci et al.CVPR 2025
- All in One Framework for Multimodal Re-Identification in the WildHe Li, Mang Ye, Ming Zhang, Bo DuCVPR 2024
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
Related papers
- World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language ModelsZiqiao Ma, Jiayi Pan, Joyce ChaiACL 2023 · 4 citations
- Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video GroundingZihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin et al.CVPR 2023
- Compositional Entailment Learning for Hyperbolic Vision-Language ModelsAvik Pal, Max van Spengler, Guido Maria D'Amely di Melendugno, Alessandro Flaborea et al.ICLR 2025
- One-Stage Visual Grounding via Semantic-Aware Feature FilterJiabo Ye, Xin Lin, Liang He, Dingbang Li et al.ACM MM 2021 · 38 citations
- Learning to Represent Image and Text with Denotation GraphBowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie et al.EMNLP 2020 · 22 citations
