Multiple Token Divergence: Measuring and Steering In-Context Computation Density
Vincent Herrmann, Eric Alcaide, Michael Wand, Jürgen Schmidhuber
摘要
Measuring the in-context computational effort of language models is a key challenge, as metrics like next-token loss fail to capture reasoning complexity. Prior methods based on latent state compressibility can be invasive and unstable. We propose Multiple Token Divergence (MTD), a simple measure of computational effort defined as the KL divergence between a model's full output distribution and that of a shallow, auxiliary prediction head. MTD can be computed directly from pre-trained models with multiple prediction heads, requiring no additional training. Building on this, we introduce Divergence Steering, a novel decoding method to control the computational character of generated text. We empirically show that MTD is more effective than prior methods at distinguishing complex tasks from simple ones. On mathematical reasoning benchmarks, MTD correlates positively with problem difficulty. Lower MTD is associated with more accurate reasoning. MTD provides a practical, lightweight tool for analyzing and steering the computational dynamics of language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz 等ICML 2024 · 被引用 286 次
- In-Context Language Learning: Architectures and AlgorithmsEkin Akyürek, Bailin Wang, Yoon Kim, Jacob AndreasICML 2024 · 被引用 91 次
- Roll the dice & look before you leap: Going beyond the creative limits of next-token predictionVaishnavh Nagarajan, Chen Henry Wu, Charles Ding, Aditi RaghunathanICML 2025
相关 Paper
- Measuring In-Context Computation Complexity via Hidden State PredictionVincent Herrmann, Róbert Csordás, Jürgen SchmidhuberICML 2025
- Don't Ignore the Tail: Decoupled Distillation Produces Top Maths Students on an Academic BudgetSayantan Dasgupta, Trevor Cohn, Tim BaldwinICML 2026
- The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE TrainingWeize Chen, Jiarui Yuan, Tailin Jin, Ning Ding 等NeurIPS 2025 · 被引用 13 次
- Auto-Regressive Next-Token Predictors are Universal LearnersEran MalachICML 2024 · 被引用 65 次
- Regress, Don't Guess: A Regression-like Loss on Number Tokens for Language ModelsJonas Zausinger, Lars Pennig, Anamarija Kozina, Sean Sdahl 等ICML 2025
