Arithmetic Sampling: Parallel Diverse Decoding for Large Language Models
Luke Vilnis, Yury Zemlyanskiy, Patrick Murray, Alexandre Tachard Passos, Sumit Sanghai
摘要
Decoding methods for large language models often trade-off between diversity of outputs and parallelism of computation. Methods such as beam search and Gumbel top-k sampling can guarantee a different output for each element of the beam, but are not easy to parallelize. Alternatively, methods such as temperature sampling and its modifications (top-k sampling, nucleus sampling, typical decoding, and others), are embarrassingly parallel, but have no guarantees about duplicate samples. We present a framework for sampling according to an arithmetic code book implicitly defined by a large language model, compatible with common sampling variations, with provable beam diversity under certain conditions, as well as being embarrassingly parallel and providing unbiased and consistent expectations from the original model. We demonstrate the effectiveness of our approach on WMT machine translation, more than halving the standard deviation when estimating expected BLEU score reward, and closing the BLEU score gap between independent sampling and beam search by up to 63%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Probabilistic Inference in Language Models via Twisted Sequential Monte CarloStephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Baker GrosseICML 2024 · 被引用 61 次
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith 等ICLR 2026 · 被引用 41 次
- Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean DiscrepancyShuhai Zhang, Yiliao Song, Jiahao Yang, Yuanqing Li 等ICLR 2024 · 被引用 19 次
- Semantic-guided Diverse Decoding for Large Language ModelWeijie Shi, Yue Cui, Yaguang Wu, Jingzhi Fang 等NeurIPS 2025 · 被引用 8 次
- p-less Sampling: A Robust Hyperparameter-Free Approach for LLM DecodingRunyan Tan, Shuang Wu, Phillip HowardICLR 2026 · 被引用 2 次
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- Determinantal Beam SearchClara Meister, Martina Forster, Ryan CotterellACL 2021
相关 Paper
- Entropy-informed Decoding: Adaptive Information-Driven BranchingBenjamin Patrick Evans, Sumitra Ganesh, Leo ArdonICML 2026
- Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM OutputsNguyen Nhat Minh, Andrew Baker, Clement Neo, Allen G. Roush 等ICLR 2025
- When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMsAmmar Khairi, Daniel D'souza, Ye Shen, Julia Kreutzer 等EMNLP 2025 · 被引用 1 次
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta 等ICLR 2024 · 被引用 31 次
- Parallel Token Prediction for Language ModelsFelix Draxler, Justus C. Will, Farrin Marouf Sofian, Theofanis Karaletsos 等ICLR 2026 · 被引用 6 次
