ModeX: Evaluator-Free Best-of-N Selection for Open-Ended Generation
Hyeong Kyu Choi, Sharon Li
Abstract
Selecting a single high-quality output from multiple stochastic generations remains a fundamental challenge for large language models (LLMs), particularly in open-ended tasks where no canonical answer exists. While Best-of-N and self-consistency methods show that aggregating multiple generations can improve performance, existing approaches typically rely on external evaluators, reward models, or exact string-match voting, limiting their applicability and efficiency. We propose Mode Extraction (ModeX), an evaluator-free Best-of-N selection framework that generalizes majority voting to open-ended text generation by identifying the modal output representing the dominant semantic consensus among generated texts. ModeX constructs a similarity graph over candidate generations and recursively applies spectral clustering to select a representative centroid, without requiring additional inference or auxiliary models. We further instantiate this selection principle as ModeX-Lite, an improved version of ModeX with early pruning for efficiency. Across open-ended tasks -- including text summarization, code generation, and mathematical reasoning -- our approaches consistently outperform standard single- and multi-path baselines, providing a computationally efficient solution for robust open-ended text generation. Code is released in https://github.com/deeplearning-wisc/ModeX.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d388717-edbc-4e44-a10e-509115bfbc45Cited by top-tier papers3
- When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent ReasoningHyeong Kyu Choi, Xiaojin (Jerry) Zhu, Sharon LiACL 2026 · 12 citations
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li et al.ACL 2026 · 8 citations
- Beyond Logits: Metastable Latent Dynamics for Sample-Efficient Best-of-N Selection in LLMsXinrong Li, Zidong Zhou, Keyu Shen, Wenhao Zhou et al.ICML 2026
Builds on14
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- Scalable Best-of-N Selection for Large Language Models via Self-CertaintyZhewei Kang, Xuandong Zhao, Dawn SongNeurIPS 2025 · 211 citations
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He et al.ICLR 2026 · 211 citations
Related papers
- MIDGARD: Self-Consistency Using Minimum Description Length for Structured Commonsense ReasoningInderjeet Nair, Lu WangACL 2024
- Lightweight reranking for language model generationsSiddhartha Jain, Xiaofei Ma, Anoop Deoras, Bing XiangACL 2024
- Do LLMs Signal When They’re Right? Evidence from Neuron AgreementKang Chen, Yaoning Wang, Kai Xiong, Zhuoka Feng et al.ICML 2026 · 8 citations
- Dynamic-Static Synergistic Selection Method for Candidate Code Solutions with Generated Test CasesRenbiao Liu, Jiang-Tian Xue, Chao-Zeng Ma, Hui Sun et al.AAAI 2026 · 2 citations
- Think in Parallel, Answer as One: Logit Averaging for Open-Ended ReasoningHaonan Wang, Chao Du, Kenji Kawaguchi, Tianyu PangICLR 2026 · 4 citations
