Combining Vision and Language Representations for Patch-based Identification of Lexico-Semantic Relations
Prince Jha, Gaël Dias, Alexis Lechervy, José G. Moreno, Anubhav Jangra, Sebastião Pais, Sriparna Saha
Abstract
Although a wide range of applications have been proposed in the field of multimodal natural language processing, very few works have been tackling multimodal relational lexical semantics. In this paper, we propose the first attempt to identify lexico-semantic relations with visual clues, which embody linguistic phenomena such as synonymy, co-hyponymy or hypernymy. While traditional methods take advantage of the paradigmatic approach or/and the distributional hypothesis, we hypothesize that visual information can supplement the textual information, relying on the apperceptum subcomponent of the semiotic textology linguistic theory. For that purpose, we automatically extend two gold-standard datasets with visual information, and develop different fusion techniques to combine textual and visual modalities following the patch-based strategy. Experimental results over the multimodal datasets show that the visual information can supplement the missing semantics of textual encodings with reliable performance improvements 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76993125-3a85-4a32-a4a2-9f8f01291c32Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Inducing Relational Knowledge from BERTZied Bouraoui, José Camacho-Collados, Steven SchockaertAAAI 2020 · 183 citations
- Multimodal Relation Extraction with Efficient Graph AlignmentChangmeng Zheng, Junhao Feng, Ze Fu, Yi Cai et al.ACM MM 2021 · 134 citations
- Multimodal Video Summarization via Time-Aware TransformersXindi Shang, Zehuan Yuan, Anran Wang, Changhu WangACM MM 2021 · 31 citations
- Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal VisionXiaoshi Wu, Hadar Averbuch-Elor, Jin Sun, Noah SnavelyICCV 2021 · 26 citations
Related papers
- Modelling Form-Meaning Systematicity with Linguistic and Visual FeaturesArie Soeteman, E. Dario Gutiérrez, Elia Bruni, Ekaterina ShutovaAAAI 2020
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu et al.ACM MM 2023 · 17 citations
- Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERFei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing et al.ACM MM 2022 · 59 citations
- A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesYunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding et al.ACL 2023 · 10 citations
- Prompt Me Up: Unleashing the Power of Alignments for Multimodal Entity and Relation ExtractionXuming Hu, Junzhe Chen, Aiwei Liu, Shiao Meng et al.ACM MM 2023 · 30 citations
