EEL: Efficiently Encoding Lattices for Reranking
Prasann Singhal, Jiacheng Xu, Xi Ye, Greg Durrett
Abstract
Standard decoding approaches for conditional text generation tasks typically search for an output hypothesis with high model probability, but this may not yield the best hypothesis according to human judgments of quality. Reranking to optimize for “downstream” metrics can more closely optimize for quality, but many metrics of interest are computed with pre-trained language models, which are slow to apply to large numbers of hypotheses. We explore an approach for reranking hypotheses by using Transformers to efficiently encode lattices of generated outputs, a method we call EEL. With a single Transformer pass over the entire lattice, we can approximately compute a contextualized representation of each token as if it were only part of a single hypothesis in isolation. We combine this approach with a new class of token-factored rerankers (TFRs) that allow for efficient extraction of high reranker-scoring hypotheses from the lattice. Empirically, our approach incurs minimal degradation error compared to the exponentially slower approach of encoding each hypothesis individually. When applying EEL with TFRs across three text generation tasks, our results show both substantial speedup compared to naive reranking and often better performance on downstream metrics than comparable approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b27d906-fed2-466e-b964-d1da84ab04fdCited by top-tier papers2
- QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine TranslationGonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei et al.NeurIPS 2024 · 23 citations
- Reranking Laws for Language Generation: A Communication-Theoretic PerspectiveAntónio Farinhas, Haau-Sing Li, André F. T. MartinsNeurIPS 2024 · 6 citations
Builds on13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Discriminative Reranking for Neural Machine TranslationAnn Lee, Michael Auli, Marc'Aurelio RanzatoACL 2021
- Language Ranker: A Lightweight Ranking framework for LLM DecodingChenheng Zhang, Tianqi Du, Jizhe Zhang, Mingqing Xiao et al.NeurIPS 2025 · 3 citations
- Fast and Accurate Deep Bidirectional Language Representations for Unsupervised LearningJoongbo Shin, Yoonhyung Lee, Seunghyun Yoon, Kyomin JungACL 2020 · 6 citations
- The Cascade Transformer: an Application for Efficient Answer Sentence SelectionLuca Soldaini, Alessandro MoschittiACL 2020 · 3 citations
- ELMER: A Non-Autoregressive Pre-trained Language Model for Efficient and Effective Text GenerationJunyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie et al.EMNLP 2022 · 11 citations
