Explaining and Improving Contrastive Decoding by Extrapolating the Probabilities of a Huge and Hypothetical LM
Haw-Shiuan Chang, Nanyun Peng, Mohit Bansal, Anil Ramakrishna, Tagyoung Chung
Abstract
Contrastive decoding (CD) (Li et al., 2023) improves the next-token distribution of a large expert language model (LM) using a small amateur LM. Although CD is applied to various LMs and domains to enhance open-ended text generation, it is still unclear why CD often works well, when it could fail, and how we can make it better. To deepen our understanding of CD, we first theoretically prove that CD could be viewed as linearly extrapolating the next-token logits from a huge and hypothetical LM. We also highlight that the linear extrapolation could make CD unable to output the most obvious answers that have already been assigned high probabilities by the amateur LM. To overcome CD's limitation, we propose a new unsupervised decoding method called Asymptotic Probability Decoding (APD). 1 APD explicitly extrapolates the probability curves from the LMs of different sizes to infer the asymptotic probabilities from an infinitely large LM without inducing more inference costs than CD. In FACTUALITYPROMPTS, an open-ended text generation benchmark, sampling using APD significantly boosts factuality in comparison to the CD sampling and its variants, and achieves state-of-the-art results for Pythia 6.9B and OPT 6.7B. Furthermore, in five commonsense QA datasets, APD is often significantly better than CD and achieves a similar effect of using a larger LLM. For example, the perplexity of APD on top of Pythia 6.9B is even lower than the perplexity of Pythia 12B in CommonsenseQA and LAMBADA. * The work was mostly done at Amazon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Generative Data Transformation: From Mixed to Unified DataJiaqing Zhang, Mingjia Yin, Hao Wang, Yuxin Tian et al.WWW 2026
- GRAD: Generalizing RAG Adaptation with DecodingYoungwon Lee, Seung-won Hwang, Zhewei Yao, Yuxiong HeACL 2026
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?Jingyuan Wang, Yankai Chen, Zhonghang Li, Chao HuangACL 2026
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
Related papers
- Contrastive Decoding: Open-ended Text Generation as OptimizationXiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang et al.ACL 2023 · 78 citations
- Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model GenerationHongxiang Zhang, Hao Chen, Muhao Chen, Tianyi ZhangEMNLP 2025 · 10 citations
- Alleviating Hallucinations in Large Language Models through Multi-Model Contrastive Decoding and Dynamic Hallucination DetectionChenyu Zhu, Yefeng Liu, Hao Zhang, Aowen Wang et al.NeurIPS 2025 · 7 citations
- Integrative Decoding: Improving Factuality via Implicit Self-consistencyYi Cheng, Xiao Liang, Yeyun Gong, Wen Xiao et al.ICLR 2025 · 1 citation
- Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive DecodingYupeng Qi, Ziyu Lyu, Lixin Cui, Lu Bai et al.ACL 2026
