Leveraging Cognitive Complexity of Texts for Contextualization in Dense Retrieval
Effrosyni Sokli, Georgios Peikos, Pranav Kasela, Gabriella Pasi
Abstract
Dense Retrieval Models (DRMs) estimate the semantic similarity between queries and documents based on their embeddings. Prior studies highlight the importance of embedding contextualization in enhancing retrieval performance. To this aim, existing approaches primarily leverage token-level information derived from query/document interactions. In this paper, we introduce a novel DRM, namely DenseC3, which leverages query/document interactions based on the full embedding representations generated by a Transformer-based model. To enhance similarity estimation, DenseC3 integrates external linguistic information about the Cognitive Complexity of texts, enriching the contextualization of embeddings. We empirically evaluate our approach across seven benchmarks and three different IR tasks to assess the impact of Cognitive Complexity-aware query and document embeddings for contextualization in dense retrieval. Results show that our approach consistently outperforms standard fine-tuning techniques on lightweight bi-encoders (e.g., BERT-based) and traditional late-interaction models (i.e., ColBERT) across all benchmarks. On larger retrieval-optimized bi-encoders like Contriever, our model achieves comparable or higher performance on four of the considered evaluation benchmarks. Our findings suggest that Cognitive Complexity-aware embeddings enhance query and document representations, improving retrieval effectiveness in DRMs. Our code is available online at: https://github.com/FaySokli/DenseC3.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17e7dff4-eeba-4016-8432-8baef3f2ea9bBuilds on5
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- SetRank: Learning a Permutation-Invariant Ranking Model for Information RetrievalLiang Pang, Jun Xu, Qingyao Ai, Yanyan Lan et al.SIGIR 2020 · 113 citations
- Mixture-of-Experts Meets Instruction Tuning: A Winning Combination for Large Language ModelsSheng Shen, Le Hou, Yanqi Zhou, Nan Du et al.ICLR 2024 · 87 citations
- COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust LearningYue Yu, Chenyan Xiong, Si Sun, Chao Zhang et al.EMNLP 2022 · 21 citations
- CODER: An efficient framework for improving retrieval through COntextual Document Embedding RerankingGeorge Zerveas, Navid Rekabsaz, Daniel Cohen, Carsten EickhoffEMNLP 2022 · 10 citations
Related papers
- Contextual Document EmbeddingsJohn Xavier Morris, Alexander M. RushICLR 2025
- Effective Contrastive Weighting for Dense Query ExpansionXiao Wang, Sean MacAvaney, Craig Macdonald, Iadh OunisACL 2023 · 2 citations
- Few-Shot Conversational Dense RetrievalShi Yu, Zhenghao Liu, Chenyan Xiong, Tao Feng et al.SIGIR 2021 · 75 citations
- Condenser: a Pre-training Architecture for Dense RetrievalLuyu Gao, Jamie CallanEMNLP 2021
- Graded Relevance Scoring of Written Essays with Dense RetrievalSalam Albatarni, Sohaila Eltanbouly, Tamer ElsayedSIGIR 2024 · 4 citations
