IR like a SIR: Sense-enhanced Information Retrieval for Multiple Languages
Rexhina Blloshmi, Tommaso Pasini, Niccolò Campolungo, Somnath Banerjee, Roberto Navigli, Gabriella Pasi
Abstract
With the advent of contextualized embeddings, attention towards neural ranking approaches for Information Retrieval increased considerably. However, two aspects have remained largely neglected: i) queries usually consist of few keywords only, which increases ambiguity and makes their contextualization harder, and ii) performing neural ranking on non-English documents is still cumbersome due to shortage of labeled datasets. In this paper we present SIR (Sense-enhanced Information Retrieval) to mitigate both problems by leveraging word sense information. At the core of our approach lies a novel multilingual query expansion mechanism based on Word Sense Disambiguation that provides sense definitions as additional semantic information for the query. Importantly, we use senses as a bridge across languages, thus allowing our model to perform considerably better than its supervised and unsupervised alternatives across French, German, Italian and Spanish languages on several CLEF benchmarks, while being trained on English Robust04 data only. We release SIR at https://github.com/SapienzaNLP/sir .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c955d2d-a76a-46a4-ab76-64f6858c197bCited by top-tier papers3
- How Much Do Encoder Models Know About Word Senses?Simone Teglia, Simone Tedeschi, Roberto NavigliACL 2025 · 1 citation
- MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense DisambiguationKaiyuan Zhang, Qian Liu, Luyang Zhang, Chaoqun Zheng et al.EMNLP 2025
- Large Scale Substitution-based Word Sense InductionMatan Eyal, Shoval Sadde, Hillel Taub-Tabib, Yoav GoldbergACL 2022
Builds on7
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Breaking Through the 80% Glass Ceiling: Raising the State of the Art in Word Sense Disambiguation by Incorporating Knowledge Graph InformationMichele Bevilacqua, Roberto NavigliACL 2020 · 145 citations
- With More Contexts Comes Better Performance: Contextualized Sense Embeddings for All-Round Word Sense DisambiguationBianca Scarlini, Tommaso Pasini, Roberto NavigliEMNLP 2020 · 95 citations
- ConSeC: Word Sense Disambiguation as Continuous Sense ComprehensionEdoardo Barba, Luigi Procopio, Roberto NavigliEMNLP 2021 · 60 citations
Related papers
- Document Translation vs. Query Translation for Cross-Lingual Information Retrieval in the Medical DomainShadi Saleh, Pavel PecinaACL 2020 · 34 citations
- Training Effective Neural CLIR by Bridging the Translation GapHamed R. Bonab, Sheikh Muhammad Sarwar, James AllanSIGIR 2020 · 35 citations
- Mind the Gap: Cross-Lingual Information Retrieval with Hierarchical Knowledge EnhancementFuwei Zhang, Zhao Zhang, Xiang Ao, Dehong Gao et al.AAAI 2022 · 26 citations
- Word Sense Disambiguation: Towards Interactive Context Exploitation from Both Word and Sense PerspectivesMing Wang, Yinglin WangACL 2021
- SensEmBERT: Context-Enhanced Sense Embeddings for Multilingual Word Sense DisambiguationBianca Scarlini, Tommaso Pasini, Roberto NavigliAAAI 2020 · 121 citations
