Forking Paths in Neural Text Generation
Eric J. Bigelow, Ari Holtzman, Hidenori Tanaka, Tomer David Ullman
摘要
Estimating uncertainty in Large Language Models (LLMs) is important for properly evaluating LLMs, and ensuring safety for users. However, prior approaches to uncertainty estimation focus on the final answer in generated text, ignoring intermediate steps that might dramatically impact the outcome. We hypothesize that there exist key forking tokens, such that re-sampling the system at those specific tokens, but not others, leads to very different outcomes. To test this empirically, we develop a novel approach to representing uncertainty dynamics across individual tokens of text generation, and applying statistical models to test our hypothesis. Our approach is highly flexible: it can be applied to any dataset and any LLM, without fine tuning or accessing model weights. We use our method to analyze LLM responses on 7 different tasks across 4 domains, spanning a wide range of typical use cases. We find many examples of forking tokens, including surprising ones such as punctuation marks, suggesting that LLMs are often just a single token away from saying something very different.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-ThoughtSiddharth Boppana, Annabel Ma, Max Loeffler, Raphaël Sarfati 等ICML 2026 · 被引用 30 次
- Thought Branches: Interpreting LLM Reasoning Requires ResamplingUzay Macar, Paul C. Bogdan, Senthooran Rajamanoharan, Neel NandaICLR 2026 · 被引用 13 次
- DISC: Dynamic Decomposition Improves LLM Inference ScalingJonathan Light, Wei Cheng, Benjamin Rivière, Yue Wu 等NeurIPS 2025 · 被引用 12 次
- TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning TasksVansh Kapoor, Aman Gupta, Hao Chen, Anurag Beniwal 等ICLR 2026 · 被引用 7 次
- The Potential of CoT for Reasoning: A Closer Look at Trace DynamicsGregor Bachmann, Yichen Jiang, Seyed-Mohsen Moosavi-Dezfooli, Moin NabiICLR 2026 · 被引用 5 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- TokUR: Token-Level Uncertainty Estimation for Large Language Model ReasoningTunyu Zhang, Haizhou Shi, Yibin Wang, Hengyi Wang 等ICLR 2026 · 被引用 19 次
- Enhancing Hallucination Detection through Noise InjectionLitian Liu, Reza Pourreza, Sunny Panchal, Apratim Bhattacharyya 等ICLR 2026 · 被引用 19 次
- Distinguishing the Knowable from the Unknowable with Language ModelsGustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak 等ICML 2024 · 被引用 44 次
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language ModelsJinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny 等ACL 2024 · 被引用 28 次
- Improving Uncertainty Estimation through Semantically Diverse Language GenerationLukas Aichberger, Kajetan Schweighofer, Mykyta Ielanskyi, Sepp HochreiterICLR 2025
