SURE or Not? Investigating Semantic Understanding in Dense Retrieval Models
Lingdi Kong, Xuanang Chen, Ben He, Le Sun
Abstract
Dense retrieval has become a core technique in applications like web search and retrievalaugmented generation. Despite their empirical success, it remains unclear whether these models truly understand semantics, and to what degree they can represent semantic consistency and distinguish subtle semantic differences. To address this gap, this paper conducts a systematic investigation by introducing SURE, a benchmark for Semantic Understanding in dense REtrieval built upon the MSMARCO, NQ, and FiQA datasets. SURE characterizes semantic understanding in dense retrieval along three dimensions: semantic precision, semantic abstraction, and semantic equivalence. We evaluate ten representative models ranging from 110M to 8B parameters, including both generalpurpose and domain-specific models. Results show that current dense retrievers struggle to distinguish fine-grained semantic differences across texts with varying information density, and to recognize semantic consistency under lexical paraphrasing. Moreover, larger models do not necessarily exhibit stronger semantic understanding, and diverse training data generally enhances semantic understanding on challenging retrieval tasks. https://github.com/ icip-cas/SURE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 565ed02e-c802-4422-9437-8a1c1d553eb4Builds on10
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin et al.SIGIR 2021 · 297 citations
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical AssessmentKun Luo, Minghao Qin, Zheng Liu, Shitao Xiao et al.EMNLP 2024 · 3 citations
Related papers
- Revisiting Single-Table Retrieval: An Open Problem Under 360° Stress TestsChenyu Yang, Ziyu Jiang, Junhao Li, Yuyu Luo et al.ICDE 2026
- Phrase Retrieval Learns Passage Retrieval, TooJinhyuk Lee, Alexander Wettig, Danqi ChenEMNLP 2021 · 1 citation
- Adversarial Retriever-Ranker for Dense Text RetrievalHang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv et al.ICLR 2022 · 137 citations
- DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search EngineYifu Qiu, Hongyu Li, Yingqi Qu, Ying Chen et al.EMNLP 2022 · 10 citations
- ConvMix: A Mixed-Criteria Data Augmentation Framework for Conversational Dense RetrievalFengran Mo, Jinghan Zhang, Yuchen Hui, Jia Ao Sun et al.AAAI 2026 · 7 citations
