Principled Zero-shot Ranking Agents with Tournament Graphs
Sheshansh Agrawal, Thien Nguyen, Douwe Kiela
Abstract
Selecting the top m from n items via expensive k-wise comparisons is central to settings ranging from LLM-based document reranking to crowdsourced evaluation and tournament design. Existing methods either rely on heuristics that discard comparison information, or exploit it at prohibitive cost. We introduce a tournament graph framework that provides a principled foundation for k-wise ranking. Our key observation is that each k-item comparison reveals an induced tournament of k 2 pairwise preferences; aggregating these into a global preference graph and computing its transitive closure yields many additional orderings without further oracle calls. We formalize when the current top-m output is certifiably determined and design a greedy query schedule that maximizes information gain towards identifying the top-m items. The framework also gracefully handles non-transitive preferences -cycles induced by real-world oracles -by collapsing them into equivalence classes that yield principled tiered rankings. Applied to LLM reranking across 14 benchmarks and 5 models, BLITZRANK achieves Pareto dominance over existing approaches: matching or exceeding accuracy while requiring 25-40% fewer tokens than comparable methods; against pairwise reranking, it achieves near-identical quality with 7× fewer tokens. Code available at https://github. com/ContextualAI/BlitzRank .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3861ca39-a6ed-4bb4-bfaf-3490b22ae186Builds on6
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan et al.EMNLP 2022 · 69 citations
- Optimal Bounds for Noisy SortingYuzhou Gu, Yinzhan XuSTOC 2023 · 11 citations
- Sample Complexity Bounds for Active Ranking from Multi-wise ComparisonsWenbo Ren, Jia Liu, Ness B. ShroffNeurIPS 2021 · 5 citations
- ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph ReconstructionYan Yu, Yilun Liu, Minggui He, Shimin Tao et al.AAAI 2026 · 2 citations
Related papers
- BracketRank: Large Language Model Document Ranking via Reasoning-based Competitive EliminationAbdelrahman Abdallah, Mohammed Ali, Bhawna Piryani, Adam JatowtACL 2026
- Beyond Sequential Reranking: Reranker-Guided Search Improves Reasoning Intensive RetrievalHaike Xu, Tong ChenICLR 2026 · 4 citations
- Investigating Non-Transitivity in LLM-as-a-JudgeYi Xu, Laura Ruis, Tim Rocktäschel, Robert KirkICML 2025
- On The Structure of Parametric Tournaments with Application to Ranking from Pairwise ComparisonsVishnu Veerathu, Arun RajkumarNeurIPS 2021 · 2 citations
- REALM: Recursive Relevance Modeling for LLM-based Document Re-RankingPinhuan Wang, Zhiqiu Xia, Chunhua Liao, Feiyi Wang et al.EMNLP 2025
