Multiple Token Divergence: Measuring and Steering In-Context Computation Density
Vincent Herrmann, Eric Alcaide, Michael Wand, Jürgen Schmidhuber
Abstract
Measuring the in-context computational effort of language models is a key challenge, as metrics like next-token loss fail to capture reasoning complexity. Prior methods based on latent state compressibility can be invasive and unstable. We propose Multiple Token Divergence (MTD), a simple measure of computational effort defined as the KL divergence between a model's full output distribution and that of a shallow, auxiliary prediction head. MTD can be computed directly from pre-trained models with multiple prediction heads, requiring no additional training. Building on this, we introduce Divergence Steering, a novel decoding method to control the computational character of generated text. We empirically show that MTD is more effective than prior methods at distinguishing complex tasks from simple ones. On mathematical reasoning benchmarks, MTD correlates positively with problem difficulty. Lower MTD is associated with more accurate reasoning. MTD provides a practical, lightweight tool for analyzing and steering the computational dynamics of language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29b33eff-3d6b-4509-bc64-a16b6ba528e3Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz et al.ICML 2024 · 286 citations
- In-Context Language Learning: Architectures and AlgorithmsEkin Akyürek, Bailin Wang, Yoon Kim, Jacob AndreasICML 2024 · 91 citations
- Roll the dice & look before you leap: Going beyond the creative limits of next-token predictionVaishnavh Nagarajan, Chen Henry Wu, Charles Ding, Aditi RaghunathanICML 2025
Related papers
- Measuring In-Context Computation Complexity via Hidden State PredictionVincent Herrmann, Róbert Csordás, Jürgen SchmidhuberICML 2025
- Don't Ignore the Tail: Decoupled Distillation Produces Top Maths Students on an Academic BudgetSayantan Dasgupta, Trevor Cohn, Tim BaldwinICML 2026
- The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE TrainingWeize Chen, Jiarui Yuan, Tailin Jin, Ning Ding et al.NeurIPS 2025 · 13 citations
- Auto-Regressive Next-Token Predictors are Universal LearnersEran MalachICML 2024 · 65 citations
- Regress, Don't Guess: A Regression-like Loss on Number Tokens for Language ModelsJonas Zausinger, Lars Pennig, Anamarija Kozina, Sean Sdahl et al.ICML 2025
