Twist Decoding: Diverse Generators Guide Each Other
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Hao Peng, Ximing Lu, Dragomir Radev, Yejin Choi, Noah A. Smith
Abstract
Many language generation models are now available for a wide range of generation tasks, including machine translation and summarization. Combining such diverse models may lead to further progress, but ensembling generation models is challenging during inference: conventional ensembling methods (e.g., shallow fusion) require that the models share vocabulary/tokenization schemes. We introduce TWIST decoding, a simple and general text generation algorithm that benefits from diverse models at inference time. Our method does not assume the vocabulary, tokenization or even generation order is shared. Our extensive evaluations on machine translation and scientific paper summarization demonstrate that TWIST decoding substantially outperforms each model decoded in isolation over various scenarios, including cases where domainspecific and general-purpose models are both available. TWIST decoding also consistently outperforms the popular reranking heuristic where output candidates from one model are rescored by another. We hope that our work will encourage researchers and practitioners to examine generation models collectively, not just independently, and to seek out models with complementary strengths to the currently available models. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d443d11-b454-45cf-975b-958104938900Cited by top-tier papers4
- T-REG: Preference Optimization with Token-Level Reward RegularizationWenxuan Zhou, Shujian Zhang, Lingxiao Zhao, Tao MengACL 2025 · 11 citations
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 9 citations
- Anchored Decoding: Provably Reducing Copyright Risk for Any Language ModelJacqueline He, Jonathan Hayase, Scott Yih, Sewoong Oh et al.ICML 2026
- Mixtures of In-Context LearnersGiwon Hong, Emile van Krieken, Edoardo Maria Ponti, Nikolay Malkin et al.ACL 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz et al.ICLR 2021 · 425 citations
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 113 citations
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 86 citations
Related papers
- Sequence Generation with Mixed RepresentationsLijun Wu, Shufang Xie, Yingce Xia, Yang Fan et al.ICML 2020 · 18 citations
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang et al.NeurIPS 2024 · 94 citations
- One Source, Two Targets: Challenges and Rewards of Dual DecodingJitao Xu, François YvonEMNLP 2021
- AdaFuse: Adaptive Ensemble Decoding for Large Language ModelsChengming Cui, Tianxin Wei, Ziyi Chen, Ruizhong Qiu et al.ACL 2026
- Transductive Ensemble Learning for Neural Machine TranslationYiren Wang, Lijun Wu, Yingce Xia, Tao Qin et al.AAAI 2020 · 29 citations
