Rethinking Reasoning in Document Ranking: Why Chain-of-Thought Falls Short
Xuan Lu, Haohang Huang, Rui Meng, Yaohui Jin, Wenjun Zeng, Xiaoyu Shen
摘要
Document reranking is a key component in information retrieval (IR), aimed at refining initial retrieval results to improve ranking quality for downstream tasks. Recent studies—motivated by large reasoning models (LRMs)—have begun incorporating explicit chain-of-thought (CoT) reasoning into LLM-based rerankers. However, the effectiveness of such reasoning for ranking tasks remains underexplored. In this work, we present the first systematic study of reasoning in reranking across both logits-based pointwise and listwise settings, under both supervised fine-tuning and reinforcement learning. Using diverse benchmarks, including reasoning-intensive datasets BRIGHT and standard IR benchmarks BEIR, we find that reasoning-augmented rerankers consistently underperform their direct counterparts that predict rankings without CoT, despite substantially higher inference costs. Our analysis reveals three core limitations: (i) in pointwise rerankers, reasoning breaks calibration and biases models toward the positive class, raising TPR but lowering TNR, which inflates false positives and degrades ranking in negative-dominant pools; (ii) in listwise rerankers, explicit reasoning improves the fit during training but leads to higher variance and fails to improve performance in both in-domain and out-of-domain evaluations, even when reinforcement learning shortens rationales; and (iii) overall, directly fine-tuned rerankers remain more stable, effective, and robust. These findings challenge the assumption that explicit reasoning is universally beneficial for reranking. We conclude by highlighting future directions, including calibration-aware scoring for pointwise rerankers and the design of concise, targeted reasoning strategies to mitigate overfitting and overthinking in listwise rerankers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Tools are under-documented: Simple Document Expansion Boosts Tool RetrievalXuan Lu, Haohang Huang, Rui Meng, Yaohui Jin 等ICLR 2026 · 被引用 16 次
- Beyond Global Similarity: Multi-Conditional Retrieval for Fine-Grained Cross-Modal UnderstandingXuan Lu, Kangle Li, Haohang Huang, Rui Meng 等CVPR 2026
它引用的顶会 Paper6
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang 等EMNLP 2023 · 被引用 182 次
- A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language ModelsShengyao Zhuang, Honglei Zhuang, Bevan Koopman, Guido ZucconSIGIR 2024 · 被引用 60 次
- ReasonRank: Empowering Passage Ranking with Strong Reasoning AbilityWenhan Liu, Xinyu Ma, Weiwei Sun, Yutao Zhu 等ACL 2026 · 被引用 43 次
- REARANK: Reasoning Re-ranking Agent via Reinforcement LearningLe Zhang, Bo Wang, Xipeng Qiu, Siva Reddy 等EMNLP 2025 · 被引用 14 次
相关 Paper
- ERank: Fusing Supervised Fine-Tuning and Reinforcement Learning for Effective and Efficient Text RerankingYuzheng Cai, Yanzhao Zhang, Dingkun Long, Mingxin Li 等AAAI 2026 · 被引用 6 次
- TFRank: Think-Free Reasoning Enables Practical Pointwise LLM RankingYongqi Fan, Xiaoyang Chen, Dezhi Ye, Jie Liu 等AAAI 2026 · 被引用 9 次
- FIRST: Faster Improved Listwise Reranking with Single Token DecodingRevanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md. Arafat Sultan 等EMNLP 2024 · 被引用 14 次
- Reason-to-Rank: Distilling Direct and Comparative Reasoning from Large Language Models for Document RerankingYuelyu Ji, Zhuochun Li, Rui Meng, Daqing HeSIGIR 2025 · 被引用 3 次
- ElicitR: Unlocking Latent Reasoning in Dense Retrievers via Generative RegularizationFengyu Cai, Iryna Gurevych, Heinz KoepplICML 2026
