Diverse Text Decoding via Iterative Reweighting
Ruiqi Shi, Sinno Jialin Pan
Abstract
Recent advances in large language models (LLMs) have led to impressive results in text generation. However, current decoding methods still lack diversity when combined with popular sampling techniques. We propose a Reweighting-based Iterative DEcoding (OverRIDE) approach that dynamically adjusts the decoding process with history responses. Our method fine-tunes auxiliary output heads iteratively on previously generated sequences to capture and suppress semantic patterns that appear in the history responses. This inference-time training process only incurs minimal loss of efficiency. We conduct extensive experiments on various tasks, including code generation, mathematical reasoning and story generation, demonstrating that OverRIDE increases output diversity while maintaining quality. We implement OverRIDE on LLM serving systems like vLLM, achieving a 6.4% throughput loss for 72B models under parallel decoding. The code is available at https://github.com/shi-rq/OverRIDE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1233320-72fa-4070-ab6a-7a7120a5a72aCited by top-tier papers2
- Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy OptimizationXingyuan Hua, Sheng Yue, Ju RenICML 2026 · 1 citation
- Large Language Models Explore by Latent DistillingYuanhao Zeng, Ao Lu, Lufei Li, Zheng Zhang et al.ICML 2026
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- Avoidance Decoding for Diverse Multi-Branch Story GenerationKyeongman Park, Nakyeong Yang, Kyomin JungEMNLP 2025 · 1 citation
- Language Ranker: A Lightweight Ranking framework for LLM DecodingChenheng Zhang, Tianqi Du, Jizhe Zhang, Mingqing Xiao et al.NeurIPS 2025 · 3 citations
- Approximately Aligned DecodingDaniel Melcer, Sujan Kumar Gonugondla, Pramuditha Perera, Haifeng Qian et al.NeurIPS 2025 · 3 citations
- Semantic-guided Diverse Decoding for Large Language ModelWeijie Shi, Yue Cui, Yaguang Wu, Jingzhi Fang et al.NeurIPS 2025 · 8 citations
- StitchLLM: Serving LLMs, One Block at a TimeBodun Hu, Shuozhe Li, Saurabh Agarwal, Myungjin Lee et al.ACL 2025
