Dimension Importance Estimation for Dense Information Retrieval
Guglielmo Faggioli, Nicola Ferro, Raffaele Perego, Nicola Tonellotto
摘要
Recent advances in Information Retrieval have shown the effectiveness of embedding queries and documents in a latent high-dimensional space to compute their similarity. While operating on such high-dimensional spaces is effective, in this paper, we hypothesize that we can improve the retrieval performance by adequately moving to a query-dependent subspace. More in detail, we formulate the Manifold Clustering (MC) Hypothesis: projecting queries and documents onto a subspace of the original representation space can improve retrieval effectiveness. To empirically validate our hypothesis, we define a novel class of Dimension IMportance Estimators (DIME). Such models aim to determine how much each dimension of a high-dimensional representation contributes to the quality of the final ranking and provide an empirical method to select a subset of dimensions where to project the query and the documents. To support our hypothesis, we propose an oracle DIME, capable of effectively selecting dimensions and almost doubling the retrieval performance. To show the practical applicability of our approach, we then propose a set of DIMEs that do not require any oracular piece of information to estimate the importance of dimensions. These estimators allow us to carry out a dimensionality selection that enables performance improvements of up to +11.5% (moving from 0.675 to 0.752 nDCG@10) compared to the baseline methods using all dimensions. Finally, we show that, with simple and realistic active feedback, such as the user's interaction with a single relevant document, we can design a highly effective DIME, allowing us to outperform the baseline by up to +0.224 nDCG@10 points (+58.6%, moving from 0.384 to 0.608).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning to Select: Query-Aware Adaptive Dimension Selection for Dense RetrievalZhanyu Wu, Richong Zhang, Zhijie NieACL 2026
- Self-Improving Sparse Retrieval Through Heuristic Representation Refinement and Representation-Focused LearningXiaojing Li, Bin Wang, Xiaochun Yang, Meng LuoAAAI 2026
它引用的顶会 Paper7
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum 等ICLR 2021 · 被引用 381 次
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo 等SIGIR 2021 · 被引用 242 次
- Emergence of Separable Manifolds in Deep Language RepresentationsJonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson 等ICML 2020 · 被引用 52 次
相关 Paper
- CoDIME: A Counterfactual Approach for Dimension Importance Estimation through Click LogsGuglielmo Faggioli, Nicola Ferro, Raffaele Perego, Nicola TonellottoSIGIR 2025 · 被引用 2 次
- Improving Document Representations by Generating Pseudo Query Embeddings for Dense RetrievalHongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang 等ACL 2021
- MA-DPR: Manifold-aware Distance Metrics for Dense Passage RetrievalYifan Liu, Qianfeng Wen, Mark Zhao, Jiazhou Liang 等EMNLP 2025 · 被引用 5 次
- Implicit Relative Labeling-Importance Aware Multi-Label Metric LearningJunxiang Mao, Yong Rui, Min-Ling ZhangAAAI 2025 · 被引用 3 次
- Towards Escaping from Class Dependency Modeling for Multi-Dimensional ClassificationTeng Huang, Bin-Bin Jia, Min-Ling ZhangICML 2025
