Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing
Le Yan, Zhen Qin, Honglei Zhuang, Rolf Jagerman, Xuanhui Wang, Michael Bendersky, Harrie Oosterhuis
Abstract
The powerful generative abilities of large language models (LLMs) show potential in generating relevance labels for search applications. Previous work has found that directly asking about relevancy, such as "How relevant is document A to query Q?", results in sub-optimal ranking. Instead, the pairwise-ranking prompting (PRP) approach produces promising ranking performance through asking about pairwise comparisons, e.g., "Is document A more relevant than document B to query Q?". Thus, while LLMs are effective at their ranking ability, this is not reflected in their relevance label generation. In this work, we propose a post-processing method to consolidate the relevance labels generated by an LLM with its powerful ranking abilities. Our method takes both LLM generated relevance labels and pairwise preferences. The labels are then altered to satisfy the pairwise preferences of the LLM, while staying as close to the original values as possible. Our experimental results indicate that our approach effectively balances label accuracy and ranking performance. Thereby, our work shows it is possible to combine both the ranking and labeling abilities of LLMs through post-processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 620e8e5c-df8e-442f-8c90-cb02ab923b3bCited by top-tier papers5
- Post Hoc Regression Refinement via Pairwise RankingsKevin Tirta Wijaya, Michael Sun, Minghao Guo, Hans-Peter Seidel et al.NeurIPS 2025 · 1 citation
- Bayesian Post Training Enhancement of Regression Models with Calibrated RankingsKevin Tirta Wijaya, Bing Hu, Hans-Peter Seidel, Wojciech Matusik et al.ICLR 2026
- Optimizing Compound Retrieval SystemsHarrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui WangSIGIR 2025
- PaSa: An LLM Agent for Comprehensive Academic Paper SearchYichen He, Guanhua Huang, Peiyuan Feng, Yuan Lin et al.ACL 2025
- Precise Zero-Shot Pointwise Ranking with LLMs through Post-Aggregated Global Context InformationKehan Long, Shasha Li, Chen Xu, Jintao Tang et al.SIGIR 2025
Builds on4
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
- Large Language Models can Accurately Predict Searcher PreferencesPaul Thomas, Seth Spielman, Nick Craswell, Bhaskar MitraSIGIR 2024 · 153 citations
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan et al.EMNLP 2022 · 69 citations
- Are Neural Rankers still Outperformed by Gradient Boosted Decision Trees?Zhen Qin, Le Yan, Honglei Zhuang, Yi Tay et al.ICLR 2021 · 41 citations
Related papers
- PRP-Graph: Pairwise Ranking Prompting to LLMs with Graph Aggregation for Effective Text Re-rankingJian Luo, Xuanang Chen, Ben He, Le SunACL 2024 · 7 citations
- A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language ModelsShengyao Zhuang, Honglei Zhuang, Bevan Koopman, Guido ZucconSIGIR 2024 · 60 citations
- Pareto Prompt OptimizationGuang Zhao, Byung-Jun Yoon, Gilchan Park, Shantenu Jha et al.ICLR 2025
- PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document RetrievalShengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin et al.EMNLP 2024 · 26 citations
- GLaPE: Gold Label-agnostic Prompt Evaluation for Large Language ModelsXuanchang Zhang, Zhuosheng Zhang, Hai ZhaoEMNLP 2024 · 3 citations
