Model-Based Minimum Bayes Risk Decoding for Text Generation
Yuu Jinnai, Tetsuro Morimura, Ukyo Honda, Kaito Ariu, Kenshi Abe
摘要
Minimum Bayes Risk (MBR) decoding has been shown to be a powerful alternative to beam search decoding in a variety of text generation tasks. MBR decoding selects a hypothesis from a pool of hypotheses that has the least expected risk under a probability model according to a given utility function. Since it is impractical to compute the expected risk exactly over all possible hypotheses, two approximations are commonly used in MBR. First, it integrates over a sampled set of hypotheses rather than over all possible hypotheses. Second, it estimates the probability of each hypothesis using a Monte Carlo estimator. While the first approximation is necessary to make it computationally feasible, the second is not essential since we typically have access to the model probability at inference time. We propose model-based MBR (MBMBR), a variant of MBR that uses the model probability itself as the estimate of the probability distribution instead of the Monte Carlo estimate. We show analytically and empirically that the model-based estimate is more promising than the Monte Carlo estimate in text generation tasks. Our experiments show that MBMBR outperforms MBR in several text generation tasks, both with encoder-decoder models and with language models. Our code is available at https://github.com/CyberAgentA ILab/model-based-mbr .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMsXueyan Li, Guinan Su, Mrinmaya Sachan, Jonas GeipingICLR 2026 · 被引用 5 次
- Noisy-Channel Minimum Bayes Risk DecodingYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeICML 2026
- Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal TransportYuu JinnaiACL 2025
它引用的顶会 Paper11
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang 等ICCV 2019 · 被引用 631 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
相关 Paper
- Case-Based Decision-Theoretic Decoding with Quality MemoriesHiroyuki Deguchi, Masaaki NagataEMNLP 2025
- Theoretical Guarantees for Minimum Bayes Risk DecodingYuki Ichihara, Yuu Jinnai, Kaito Ariu, Tetsuro Morimura 等ACL 2025
- Uncertainty-Aware Decoding with Minimum Bayes RiskNico Daheim, Clara Meister, Thomas Möllenhoff, Iryna GurevychICLR 2025
- Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion AlgorithmsFiras Trabelsi, David Vilar, Mara Finkelstein, Markus FreitagNeurIPS 2024 · 被引用 18 次
- Sampling-Based Approximations to Minimum Bayes Risk Decoding for Neural Machine TranslationBryan Eikema, Wilker AzizEMNLP 2022 · 被引用 10 次
