An Empirical Study of Translation Hypothesis Ensembling with Large Language Models
António Farinhas, José Guilherme Camargo de Souza, André F. T. Martins
摘要
Large language models (LLMs) are becoming a one-fits-many solution, but they sometimes hallucinate or produce unreliable output. In this paper, we investigate how hypothesis ensembling can improve the quality of the generated text for the specific problem of LLM-based machine translation. We experiment with several techniques for ensembling hypotheses produced by LLMs such as ChatGPT, LLaMA, and Alpaca. We provide a comprehensive study along multiple dimensions, including the method to generate hypotheses (multiple prompts, temperaturebased sampling, and beam search) and the strategy to produce the final translation (instructionbased, quality-based reranking, and minimum Bayes risk (MBR) decoding). Our results show that MBR decoding is a very effective method, that translation quality can be improved using a small number of samples, and that instruction tuning has a strong impact on the relation between the diversity of the hypotheses and the sampling temperature. Our code is available at https://github.com/deep-spin/ translation-hypothesis-ensembling .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine TranslationGonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei 等NeurIPS 2024 · 被引用 23 次
- Cool-Fusion: Fuse Large Language Models without TrainingCong Liu, Xiaojun Quan, Yan Pan, Weigang Wu 等ACL 2025 · 被引用 12 次
- Model-Based Minimum Bayes Risk Decoding for Text GenerationYuu Jinnai, Tetsuro Morimura, Ukyo Honda, Kaito Ariu 等ICML 2024 · 被引用 9 次
- Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality EstimationEmmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. MartinsACL 2025 · 被引用 8 次
- CoT-based Synthesizer: Enhancing LLM Performance through Answer SynthesisBohan Zhang, Xiaokang Zhang, Jing Zhang, Jifan Yu 等ACL 2025 · 被引用 8 次
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok 等CHI 2021 · 被引用 713 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Improving Minimum Bayes Risk Decoding with Multi-PromptDavid Heineman, Yao Dou, Wei XuEMNLP 2024 · 被引用 1 次
- Self-Alignment with Instruction BacktranslationXian Li, Ping Yu, Chunting Zhou, Timo Schick 等ICLR 2024 · 被引用 174 次
- Better Instruction-Following Through Minimum Bayes RiskIan Wu, Patrick Fernandes, Amanda Bertsch, Seungone Kim 等ICLR 2025
- Don't Rank, Combine! Combining Machine Translation Hypotheses Using Quality EstimationGiorgos Vernikos, Andrei Popescu-BelisACL 2024
- RLAE: Reinforcement Learning-Assisted Ensemble for LLMsYuqian Fu, Yuanheng Zhu, Jiajun Chai, Guojun Yin 等EMNLP 2025 · 被引用 1 次
