Beyond Logits: Metastable Latent Dynamics for Sample-Efficient Best-of-N Selection in LLMs
Xinrong Li, Zidong Zhou, Keyu Shen, Wenhao Zhou, Shangqi Guo
Abstract
Best-of-N selection improves reasoning in large language models (LLMs) by allocating test-time compute to sample candidate trajectories, but it relies on reliable verification. Widely used proxies have complementary failure modes: logit-confidence signals can suffer from calibration collapse, where confidence becomes misaligned with correctness, while sample agreement can be costly or brittle. Instead, we analyze the model's latent dynamics during inference. Motivated by metastable dynamics in cognitive systems, we introduce Latent Velocity Entropy (LVE), a training-free metric that quantifies the temporal concentration of internal representation updates. Experiments on four reasoning benchmarks, AIME25, GPQA, MATH500, and BRUMO25, show that LVE-based selection remains informative when logit-confidence baselines fail, with the clearest gains on reasoning-dense mathematical tasks and more modest gains on knowledge-heavy GPQA. On MATH500, LogNorm-LVE reaches 92.2% Pass@1 at and nearly matches 10-sample majority voting with only 3 samples. Cross-model, bootstrap, and length-controlled analyses indicate that LVE is not reducible to sequence length, while also clarifying the task-dependent limits of the signal. Code is available at https://github.com/ZopenDP/LatentVelocityEntropy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Scalable Best-of-N Selection for Large Language Models via Self-CertaintyZhewei Kang, Xuandong Zhao, Dawn SongNeurIPS 2025 · 211 citations
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian et al.ICLR 2026 · 171 citations
Related papers
- ANCHOR: Taming Entropy Dynamics for Stable and Efficient Reasoning of Large Language ModelsCong Qin, Jiaye Lin, Xiaoliang Fu, Yangyi Fang et al.KDD 2026
- I²B-LPO: Latent Policy Optimization via Iterative Information BottleneckHuilin Deng, Hongchen Luo, Yue Zhu, Long Li et al.ACL 2026
- Dissecting Failure Dynamics in Large Language Model ReasoningWei Zhu, Jian Zhang, Lixing Yu, Kun Yue et al.ACL 2026 · 2 citations
- Tracing the Traces: Latent Temporal Signals for Efficient and Accurate ReasoningMartina G. Vilas, Safoora Yousefi, Besmira Nushi, Eric Horvitz et al.ICLR 2026 · 12 citations
- Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning ModelsGuanxu Chen, Yafu Li, Yuxian Jiang, Chen Qian et al.ICLR 2026 · 3 citations
