A Thorough Examination of Decoding Methods in the Era of LLMs
Chufan Shi, Haoran Yang, Deng Cai, Zhisong Zhang, Yifan Wang, Yujiu Yang, Wai Lam
摘要
Decoding methods play an indispensable role in converting language models from next-token predictors into practical task solvers. Prior research on decoding methods, primarily focusing on task-specific models, may not extend to the current era of general-purpose large language models (LLMs). Moreover, the recent influx of decoding strategies has further complicated this landscape. This paper provides a comprehensive and multifaceted analysis of various decoding methods within the context of LLMs, evaluating their performance, robustness to hyperparameter changes, and decoding speeds across a wide range of tasks, models, and deployment environments. Our findings reveal that decoding method performance is notably task-dependent and influenced by factors such as alignment, model size, and quantization. Intriguingly, sensitivity analysis exposes that certain methods achieve superior performance at the cost of extensive hyperparameter tuning, highlighting the trade-off between attaining optimal results and the practicality of implementation in varying contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt OptimizersQingyan Guo, Rui Wang, Junliang Guo, Bei Li 等ICLR 2024 · 被引用 257 次
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman 等EMNLP 2024 · 被引用 47 次
- Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-ContrastChufan Shi, Cheng Yang, Xinyu Zhu, Jiahao Wang 等NeurIPS 2024 · 被引用 27 次
- SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language ModelsJianyi Zhang, Da-Cheng Juan, Cyrus Rashtchian, Chun-Sung Ferng 等NeurIPS 2024 · 被引用 22 次
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and GenerationLing Yang, Xinchen Zhang, Ye Tian, Shiyi Zhang 等NeurIPS 2025 · 被引用 16 次
它引用的顶会 Paper14
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim 等ICLR 2024 · 被引用 354 次
相关 Paper
- Guiding LLMs The Right Way: Fast, Non-Invasive Constrained GenerationLuca Beurer-Kellner, Marc Fischer, Martin T. VechevICML 2024 · 被引用 93 次
- Does quantization affect models' performance on long-context tasks?Anmol Mekala, Anirudh Atmakuru, Yixiao Song, Marzena Karpinska 等EMNLP 2025
- Adaptive Draft-Verification for Efficient Large Language Model DecodingXukun Liu, Bowen Lei, Ruqi Zhang, Dongkuan XuAAAI 2025 · 被引用 9 次
- A Theoretical Perspective for Speculative Decoding AlgorithmMing Yin, Minshuo Chen, Kaixuan Huang, Mengdi WangNeurIPS 2024 · 被引用 36 次
- OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMsShaoyuan Chen, Zhixuan Chen, Dawei Yang, Zhihang Yuan 等AAAI 2026
