Rethinking Reasoning in Document Ranking: Why Chain-of-Thought Falls Short
Xuan Lu, Haohang Huang, Rui Meng, Yaohui Jin, Wenjun Zeng, Xiaoyu Shen
Abstract
Document reranking is a key component in information retrieval (IR), aimed at refining initial retrieval results to improve ranking quality for downstream tasks. Recent studies—motivated by large reasoning models (LRMs)—have begun incorporating explicit chain-of-thought (CoT) reasoning into LLM-based rerankers. However, the effectiveness of such reasoning for ranking tasks remains underexplored. In this work, we present the first systematic study of reasoning in reranking across both logits-based pointwise and listwise settings, under both supervised fine-tuning and reinforcement learning. Using diverse benchmarks, including reasoning-intensive datasets BRIGHT and standard IR benchmarks BEIR, we find that reasoning-augmented rerankers consistently underperform their direct counterparts that predict rankings without CoT, despite substantially higher inference costs. Our analysis reveals three core limitations: (i) in pointwise rerankers, reasoning breaks calibration and biases models toward the positive class, raising TPR but lowering TNR, which inflates false positives and degrades ranking in negative-dominant pools; (ii) in listwise rerankers, explicit reasoning improves the fit during training but leads to higher variance and fails to improve performance in both in-domain and out-of-domain evaluations, even when reinforcement learning shortens rationales; and (iii) overall, directly fine-tuned rerankers remain more stable, effective, and robust. These findings challenge the assumption that explicit reasoning is universally beneficial for reranking. We conclude by highlighting future directions, including calibration-aware scoring for pointwise rerankers and the design of concise, targeted reasoning strategies to mitigate overfitting and overthinking in listwise rerankers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 440473c1-d448-415b-b50a-26fac9bd4e38Cited by top-tier papers2
- Tools are under-documented: Simple Document Expansion Boosts Tool RetrievalXuan Lu, Haohang Huang, Rui Meng, Yaohui Jin et al.ICLR 2026 · 16 citations
- Beyond Global Similarity: Multi-Conditional Retrieval for Fine-Grained Cross-Modal UnderstandingXuan Lu, Kangle Li, Haohang Huang, Rui Meng et al.CVPR 2026
Builds on6
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
- A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language ModelsShengyao Zhuang, Honglei Zhuang, Bevan Koopman, Guido ZucconSIGIR 2024 · 60 citations
- ReasonRank: Empowering Passage Ranking with Strong Reasoning AbilityWenhan Liu, Xinyu Ma, Weiwei Sun, Yutao Zhu et al.ACL 2026 · 43 citations
- REARANK: Reasoning Re-ranking Agent via Reinforcement LearningLe Zhang, Bo Wang, Xipeng Qiu, Siva Reddy et al.EMNLP 2025 · 14 citations
Related papers
- ERank: Fusing Supervised Fine-Tuning and Reinforcement Learning for Effective and Efficient Text RerankingYuzheng Cai, Yanzhao Zhang, Dingkun Long, Mingxin Li et al.AAAI 2026 · 6 citations
- TFRank: Think-Free Reasoning Enables Practical Pointwise LLM RankingYongqi Fan, Xiaoyang Chen, Dezhi Ye, Jie Liu et al.AAAI 2026 · 9 citations
- FIRST: Faster Improved Listwise Reranking with Single Token DecodingRevanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md. Arafat Sultan et al.EMNLP 2024 · 14 citations
- Reason-to-Rank: Distilling Direct and Comparative Reasoning from Large Language Models for Document RerankingYuelyu Ji, Zhuochun Li, Rui Meng, Daqing HeSIGIR 2025 · 3 citations
- ElicitR: Unlocking Latent Reasoning in Dense Retrievers via Generative RegularizationFengyu Cai, Iryna Gurevych, Heinz KoepplICML 2026
