Large Language Models and Multimodal Retrieval for Visual Word Sense Disambiguation
Anastasia Kritharoula, Maria Lymperaiou, Giorgos Stamou
Abstract
Visual Word Sense Disambiguation (VWSD) is a novel challenging task with the goal of retrieving an image among a set of candidates, which better represents the meaning of an ambiguous word within a given context. In this paper, we make a substantial step towards unveiling this interesting task by applying a varying set of approaches. Since VWSD is primarily a text-image retrieval task, we explore the latest transformer-based methods for multimodal retrieval. Additionally, we utilize Large Language Models (LLMs) as knowledge bases to enhance the given phrases and resolve ambiguity related to the target word. We also study VWSD as a unimodal problem by converting to text-to-text and image-to-image retrieval, as well as question-answering (QA), to fully explore the capabilities of relevant models. To tap into the implicit knowledge of LLMs, we experiment with Chain-of-Thought (CoT) prompting to guide explainable answer generation. On top of all, we train a learn to rank (LTR) model in order to combine our different modules, achieving competitive ranking results. Extensive experiments on VWSD demonstrate valuable insights to effectively drive future directions. CLIP CLIP-L ALIGN BLIPC BLIP-LC BLIPF BLIP-LF acc. MRR acc. MRR acc. MRR acc. MRR acc. MRR acc. MRR acc. MRR With penalty
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MSTAR: Box-free Multi-query Scene Text Retrieval with Attention RecyclingLiang Yin, Xudong Xie, Zhang Li, Xiang Bai et al.NeurIPS 2025 · 2 citations
- MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense DisambiguationKaiyuan Zhang, Qian Liu, Luyang Zhang, Chaoqun Zheng et al.EMNLP 2025
- PolCLIP: A Unified Image-Text Word Sense Disambiguation Model via Generating Multimodal Complementary RepresentationsQihao Yang, Yong Li, Xuelin Wang, Fu Lee Wang et al.ACL 2024
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss InformationSunjae Kwon, Rishabh Garodia, Minhwa Lee, Zhichao Yang et al.ACL 2023 · 3 citations
- CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image RetrievalZelong Sun, Dong Jing, Zhiwu LuICCV 2025 · 5 citations
- WSDPO: A Generative Word Sense Disambiguation Framework with Chain-of-Thought and Preference OptimizationKunpeng Kang, Shuaimin Li, Kaiyuan Zhang, Luyang Zhang et al.ACL 2026
- Interleaved-Modal Chain-of-ThoughtJun Gao, Yongqi Li, Ziqiang Cao, Wenjie LiCVPR 2025
- Modality-Aware Integration with Large Language Models for Knowledge-Based Visual Question AnsweringJunnan Dong, Qinggang Zhang, Huachi Zhou, Daochen Zha et al.ACL 2024 · 11 citations
