Learning grounded word meaning representations on similarity graphs
Mariella Dimiccoli, Herwig Wendt, Pau Batlle Franch
Abstract
This paper introduces a novel approach to learn visually grounded meaning representations of words as low-dimensional node embeddings on an underlying graph hierarchy. The lower level of the hierarchy models modality-specific word representations through dedicated but communicating graphs, while the higher level puts these representations together on a single graph to learn a representation jointly from both modalities. The topology of each graph models similarity relations among words, and is estimated jointly with the graph embedding. The assumption underlying this model is that words sharing similar meaning correspond to communities in an underlying similarity graph in a lowdimensional space. We named this model Hierarchical Multi-Modal Similarity Graph Embedding (HM-SGE). Experimental results validate the ability of HM-SGE to simulate human similarity judgements and concept categorization, outperforming the state of the art. 1 * Work done during an internship at the IRI (CSIC-UPC).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa5fc8ca-2f7f-4ede-a8a4-20af9084a915Builds on3
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Embedding Words in Non-Vector Space with Unsupervised Graph LearningMax Ryabinin, Sergei Popov, Liudmila Prokhorenkova, Elena VoitaEMNLP 2020 · 1 citation
Related papers
- Modelling Form-Meaning Systematicity with Linguistic and Visual FeaturesArie Soeteman, E. Dario Gutiérrez, Elia Bruni, Ekaterina ShutovaAAAI 2020
- Learning Cross-Modal Context Graph for Visual GroundingYongfei Liu, Bo Wan, Xiaodan Zhu, Xuming HeAAAI 2020 · 100 citations
- Visual-Semantic Graph Matching for Visual GroundingChenchen Jing, Yuwei Wu, Mingtao Pei, Yao Hu et al.ACM MM 2020 · 35 citations
- Enhancing Semi-Supervised Learning with Cross-Modal KnowledgeHui Zhu, Yongchun Lü, Hongbin Wang, Xunyi Zhou et al.ACM MM 2022 · 4 citations
- ViCo: Word Embeddings From Visual Co-OccurrencesTanmay Gupta, Alexander G. Schwing, Derek HoiemICCV 2019 · 26 citations
