D-Models and E-Models: Diversity-Stability Trade-offs in the Sampling Behavior of Large Language Models
Jia Gu, Liang Pang, Huawei Shen, Xueqi Cheng
摘要
The predictive probability of the next token (P_token) in large language models (LLMs) is inextricably linked to the probability of relevance for the next piece of information, the purchase probability of the next product, and the execution probability of the next action-all of which fall under the scope of the task-level target distribution (P_task). While LLMs are known to generate samples that approximate real-world distributions, whether their fine-grained sampling probabilities faithfully align with task requirements remains an open question. Through controlled distribution-sampling simulations, we uncover a striking dichotomy in LLM behavior, distinguishing two model types: D-models (e.g. Qwen-2.5), whose P_token exhibits large step-to-step variability and poor alignment with P_task; and E-models (e.g. Mistral-Small), whose P_token is more stable and better aligned with P_task. We further evaluate these two model types in downstream tasks such as code generation and recommendation, revealing systematic trade-offs between diversity and stability that shape task outcomes. Finally, we analyze the internal properties of both model families to probe their underlying mechanisms. These findings offer foundational insights into the probabilistic sampling behavior of LLMs and provide practical guidance on when to favor D- versus E-models. For web-scale applications, including recommendation, search, and conversational agents, our results inform model selection and configuration to balance diversity with reliability under real-world uncertainty, providing a better level of interpretation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- ThinkSum: Probabilistic reasoning over sets using large language modelsBatu Ozturkler, Nikolay Malkin, Zhen Wang, Nebojsa JojicACL 2023 · 被引用 10 次
- Every Answer Matters: Evaluating Commonsense with Probabilistic MeasuresQi Cheng, Michael Boratko, Pranay Kumar Yelugam, Tim O'Gorman 等ACL 2024 · 被引用 2 次
- Fourier Head: Helping Large Language Models Learn Complex Probability DistributionsNate Gillman, Daksh Aggarwal, Michael Freeman, Chen SunICLR 2025
- Enough Coin Flips Can Make LLMs Act BayesianRitwik Gupta, Rodolfo Corona, Jiaxin Ge, Eric Wang 等ACL 2025
- B-score: Detecting biases in large language models using response historyAn Vo, Mohammad Reza Taesiri, Daeyoung Kim, Anh Totti NguyenICML 2025
相关 Paper
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-FutureYidong Wang, Xin Wang, Cunxiang Wang, Junfeng Fang 等ICML 2026 · 被引用 3 次
- Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical DistributionsMinda Zhao, Yilun Du, Mengyu WangACL 2026 · 被引用 2 次
- Exploring Precision and Recall to assess the quality and diversity of LLMsFlorian Le Bronnec, Alexandre Verine, Benjamin Négrevergne, Yann Chevaleyre 等ACL 2024 · 被引用 11 次
- A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?Ibrahim Alabdulmohsin, Andreas Peter SteinerICML 2025
- Calibrating LLM Confidence by Probing Perturbed Representation StabilityReza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur 等EMNLP 2025 · 被引用 1 次
