Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid Distributions
Liuxuan Jiao, Chen Gao, Yiqian Yang, Chenliang Zhou, YiXian Huang, Xinlei Chen, Yong Li
摘要
Accurate modeling and control of response length is essential for optimizing large language model (LLM) deployment, impacting computational efficiency, user experience, and system reliability. We develop a statistical framework based on extreme value theory, analyzing 14,301 GPT-4o responses across temperature settings and prompting strategies, with cross-validation on Qwen and DeepSeek architectures. Our analysis reveals that response lengths follow Weibull-type generalized extreme value (GEV) distributions, exhibiting heavier tails under stochastic generation conditions. The key contributions include: (1) a novel GEV-generalized Pareto (GPD) hybrid model that achieves superior tail fit (R 2 CDF = 0.9993 vs standalone GEV's 0.998) while preserving architectural generalizability; (2) quantitative characterization of prompt anchoring effects, showing reduced dispersion but increased outlier propensity under randomization; and (3) identification of temperaturedependent response patterns that remain consistent across architectures, where higher temperatures amplify length variability while maintaining the underlying extreme-value mechanisms. The proposed hybrid model's adaptive threshold selection enables precise verbosity control in production systems, regardless of the specific LLM architecture employed. These findings provide both theoretical insights into LLM generation patterns and practical tools for response length optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz 等NeurIPS 2022 · 被引用 267 次
- Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference PipelineZangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo 等NeurIPS 2023 · 被引用 159 次
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsMosh Levy, Alon Jacoby, Yoav GoldbergACL 2024 · 被引用 77 次
- LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed ProblemsPengkun Wang, Zhe Zhao, Haibin Wen, Fanfu Wang 等NeurIPS 2024 · 被引用 26 次
- The Devil is in the Tails: How Long-Tailed Code Distributions Impact Large Language ModelsXin Zhou, Kisub Kim, Bowen Xu, Jiakun Liu 等ASE 2023 · 被引用 10 次
相关 Paper
- Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language ModelsThomas P. Zollo, Todd Morrill, Zhun Deng, Jake Snell 等ICLR 2024 · 被引用 14 次
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang 等ACL 2026 · 被引用 8 次
- A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?Ibrahim Alabdulmohsin, Andreas Peter SteinerICML 2025
- LLMTM: Benchmarking and Optimizing LLMs for Temporal Motif Analysis in Dynamic GraphsBing Hao, Minglai Shao, Zengyi Wo, Yunlong Chu 等AAAI 2026
- Following Length Constraints in InstructionsWeizhe Yuan, Ilia Kulikov, Ping Yu, Kyunghyun Cho 等EMNLP 2025 · 被引用 1 次
