Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss Information
Sunjae Kwon, Rishabh Garodia, Minhwa Lee, Zhichao Yang, Hong Yu
Abstract
Visual Word Sense Disambiguation (VWSD) is a task to find the image that most accurately depicts the correct sense of the target word for the given context. Previously, image-text matching models often suffered from recognizing polysemous words. This paper introduces an unsupervised VWSD approach that uses gloss information of an external lexical knowledge-base, especially the sense definitions. Specifically, we suggest employing Bayesian inference to incorporate the sense definitions when sense information of the answer is not provided. In addition, to ameliorate the out-of-dictionary (OOD) issue, we propose a context-aware definition generation with GPT-3. Experimental results show that the VWSD performance significantly increased with our Bayesian inference-based approach. In addition, our context-aware definition generation achieved prominent performance improvement in OOD examples exhibiting better performance than the existing definition generation method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense DisambiguationKaiyuan Zhang, Qian Liu, Luyang Zhang, Chaoqun Zheng et al.EMNLP 2025
- PolCLIP: A Unified Image-Text Word Sense Disambiguation Model via Generating Multimodal Complementary RepresentationsQihao Yang, Yong Li, Xuelin Wang, Fu Lee Wang et al.ACL 2024
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
- ConSeC: Word Sense Disambiguation as Continuous Sense ComprehensionEdoardo Barba, Luigi Procopio, Roberto NavigliEMNLP 2021 · 60 citations
Related papers
- Moving Down the Long Tail of Word Sense Disambiguation with Gloss Informed Bi-encodersTerra Blevins, Luke ZettlemoyerACL 2020 · 19 citations
- Large Language Models and Multimodal Retrieval for Visual Word Sense DisambiguationAnastasia Kritharoula, Maria Lymperaiou, Giorgos StamouEMNLP 2023 · 3 citations
- Word Sense Disambiguation by Refining Target Word EmbeddingXuefeng Zhang, Richong Zhang, Xiaoyang Li, Fanshuang Kong et al.WWW 2023 · 5 citations
- Connect-the-Dots: Bridging Semantics between Words and Definitions via Aligning Word Sense InventoriesWenlin Yao, Xiaoman Pan, Lifeng Jin, Jianshu Chen et al.EMNLP 2021 · 4 citations
- Do Large Language Models Understand Word Senses?Domenico Meconi, Simone Stirpe, Federico Martelli, Leonardo Lavalle et al.EMNLP 2025 · 7 citations
