Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid Distributions
Liuxuan Jiao, Chen Gao, Yiqian Yang, Chenliang Zhou, YiXian Huang, Xinlei Chen, Yong Li
Abstract
Accurate modeling and control of response length is essential for optimizing large language model (LLM) deployment, impacting computational efficiency, user experience, and system reliability. We develop a statistical framework based on extreme value theory, analyzing 14,301 GPT-4o responses across temperature settings and prompting strategies, with cross-validation on Qwen and DeepSeek architectures. Our analysis reveals that response lengths follow Weibull-type generalized extreme value (GEV) distributions, exhibiting heavier tails under stochastic generation conditions. The key contributions include: (1) a novel GEV-generalized Pareto (GPD) hybrid model that achieves superior tail fit (R 2 CDF = 0.9993 vs standalone GEV's 0.998) while preserving architectural generalizability; (2) quantitative characterization of prompt anchoring effects, showing reduced dispersion but increased outlier propensity under randomization; and (3) identification of temperaturedependent response patterns that remain consistent across architectures, where higher temperatures amplify length variability while maintaining the underlying extreme-value mechanisms. The proposed hybrid model's adaptive threshold selection enables precise verbosity control in production systems, regardless of the specific LLM architecture employed. These findings provide both theoretical insights into LLM generation patterns and practical tools for response length optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75cd91c7-c8a5-49ac-89fd-18d9dcd74685Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
- Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference PipelineZangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo et al.NeurIPS 2023 · 159 citations
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsMosh Levy, Alon Jacoby, Yoav GoldbergACL 2024 · 77 citations
- LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed ProblemsPengkun Wang, Zhe Zhao, Haibin Wen, Fanfu Wang et al.NeurIPS 2024 · 26 citations
- The Devil is in the Tails: How Long-Tailed Code Distributions Impact Large Language ModelsXin Zhou, Kisub Kim, Bowen Xu, Jiakun Liu et al.ASE 2023 · 10 citations
Related papers
- Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language ModelsThomas P. Zollo, Todd Morrill, Zhun Deng, Jake Snell et al.ICLR 2024 · 14 citations
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang et al.ACL 2026 · 8 citations
- A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?Ibrahim Alabdulmohsin, Andreas Peter SteinerICML 2025
- LLMTM: Benchmarking and Optimizing LLMs for Temporal Motif Analysis in Dynamic GraphsBing Hao, Minglai Shao, Zengyi Wo, Yunlong Chu et al.AAAI 2026
- Following Length Constraints in InstructionsWeizhe Yuan, Ilia Kulikov, Ping Yu, Kyunghyun Cho et al.EMNLP 2025 · 1 citation
