Language Model Cascades: Token-Level Uncertainty And Beyond
Neha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar
摘要
Recent advances in language models (LMs) have led to significant improvements in quality on complex NLP tasks, but at the expense of increased inference costs. Cascading offers a simple strategy to achieve more favorable cost-quality tradeoffs: here, a small model is invoked for most"easy"instances, while a few"hard"instances are deferred to the large model. While the principles underpinning cascading are well-studied for classification tasks - with deferral based on predicted class uncertainty favored theoretically and practically - a similar understanding is lacking for generative LM tasks. In this work, we initiate a systematic study of deferral rules for LM cascades. We begin by examining the natural extension of predicted class uncertainty to generative LM tasks, namely, the predicted sequence uncertainty. We show that this measure suffers from the length bias problem, either over- or under-emphasizing outputs based on their lengths. This is because LMs produce a sequence of uncertainty values, one for each output token; and moreover, the number of output tokens is variable across examples. To mitigate this issue, we propose to exploit the richer token-level uncertainty information implicit in generative LMs. We argue that naive predicted sequence uncertainty corresponds to a simple aggregation of these uncertainties. By contrast, we show that incorporating token-level uncertainty through learned post-hoc deferral rules can significantly outperform such simple aggregation strategies, via experiments on a range of natural language benchmarks with FLAN-T5 models. We further show that incorporating embeddings from the smaller model and intermediate layers of the larger model can give an additional boost in the overall cost-quality tradeoff.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyMehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld 等ICLR 2026 · 被引用 116 次
- RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language ModelsShuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok 等NeurIPS 2024 · 被引用 113 次
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja 等ICLR 2026 · 被引用 99 次
- Reasoning Models Better Express Their ConfidenceDongkeun Yoon, Seungone Kim, Sohee Yang, Sunkyoung Kim 等NeurIPS 2025 · 被引用 77 次
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for ReasoningXuechen Zhang, Zijian Huang, Yingcong Li, Chenshun Ni 等NeurIPS 2025 · 被引用 66 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
相关 Paper
- Faster Cascades via Speculative DecodingHarikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Seungyeon Kim 等ICLR 2025
- Online Cascade Learning for Efficient Inference over StreamsLunyiu Nie, Zhimin Ding, Erdong Hu, Christopher M. Jermaine 等ICML 2024 · 被引用 20 次
- Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware DeferralAntónio Farinhas, Nuno Miguel Guerreiro, Sweta Agrawal, Ricardo Rei 等EMNLP 2025 · 被引用 2 次
- Gatekeeper: Improving Model Cascades Through Confidence TuningStephan Rabanser, Nathalie Rauschmayr, Achin Kulshrestha, Petra Poklukar 等NeurIPS 2025 · 被引用 10 次
- Model Cascading: Towards Jointly Improving Efficiency and Accuracy of NLP SystemsNeeraj Varshney, Chitta BaralEMNLP 2022 · 被引用 11 次
