Forking Paths in Neural Text Generation
Eric J. Bigelow, Ari Holtzman, Hidenori Tanaka, Tomer David Ullman
Abstract
Estimating uncertainty in Large Language Models (LLMs) is important for properly evaluating LLMs, and ensuring safety for users. However, prior approaches to uncertainty estimation focus on the final answer in generated text, ignoring intermediate steps that might dramatically impact the outcome. We hypothesize that there exist key forking tokens, such that re-sampling the system at those specific tokens, but not others, leads to very different outcomes. To test this empirically, we develop a novel approach to representing uncertainty dynamics across individual tokens of text generation, and applying statistical models to test our hypothesis. Our approach is highly flexible: it can be applied to any dataset and any LLM, without fine tuning or accessing model weights. We use our method to analyze LLM responses on 7 different tasks across 4 domains, spanning a wide range of typical use cases. We find many examples of forking tokens, including surprising ones such as punctuation marks, suggesting that LLMs are often just a single token away from saying something very different.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-ThoughtSiddharth Boppana, Annabel Ma, Max Loeffler, Raphaël Sarfati et al.ICML 2026 · 30 citations
- Thought Branches: Interpreting LLM Reasoning Requires ResamplingUzay Macar, Paul C. Bogdan, Senthooran Rajamanoharan, Neel NandaICLR 2026 · 13 citations
- DISC: Dynamic Decomposition Improves LLM Inference ScalingJonathan Light, Wei Cheng, Benjamin Rivière, Yue Wu et al.NeurIPS 2025 · 12 citations
- TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning TasksVansh Kapoor, Aman Gupta, Hao Chen, Anurag Beniwal et al.ICLR 2026 · 7 citations
- The Potential of CoT for Reasoning: A Closer Look at Trace DynamicsGregor Bachmann, Yichen Jiang, Seyed-Mohsen Moosavi-Dezfooli, Moin NabiICLR 2026 · 5 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- TokUR: Token-Level Uncertainty Estimation for Large Language Model ReasoningTunyu Zhang, Haizhou Shi, Yibin Wang, Hengyi Wang et al.ICLR 2026 · 19 citations
- Enhancing Hallucination Detection through Noise InjectionLitian Liu, Reza Pourreza, Sunny Panchal, Apratim Bhattacharyya et al.ICLR 2026 · 19 citations
- Distinguishing the Knowable from the Unknowable with Language ModelsGustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak et al.ICML 2024 · 44 citations
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language ModelsJinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny et al.ACL 2024 · 28 citations
- Improving Uncertainty Estimation through Semantically Diverse Language GenerationLukas Aichberger, Kajetan Schweighofer, Mykyta Ielanskyi, Sepp HochreiterICLR 2025
