What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation
Shaomu Tan, Dawei Zhu, Ke Tran, Michael J. Denkowski, Sony Trenous, Leonardo F. R. Ribeiro, Bill Byrne, Felix Hieber
摘要
Iterative self-refinement is a simple inference-time strategy for machine translation: an LLM revises its own translation over multiple inference-time passes. Yet document-scale refinement remains poorly understood: 1) which pipelines work best, 2) what quality dimensions improve, and 3) how refiners behave. In this paper, we present a systematic study of document-level literary translation, covering nine LLMs and seven language pairs. Across nine translation-refinement granularity combinations and five refinement strategies, we find a robust recipe: document-level MT followed by segment-level refinement yields strong and stable improvements. In contrast, document-level refinement often makes fewer edits and leads to smaller or less reliable gains. Beyond granularity, A simple general refinement prompt consistently outperforms error-specific prompting and evaluate-then-refine schemes. Our large-scale human evaluation shows that refinement gains come primarily from fluency, style, and terminology, with limited and less consistent improvements in adequacy. Experiments varying model strength reveal refinement projects outputs toward the refiner's distribution rather than performing targeted error repair. These findings clarify the mechanisms and limitations of current refinement approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang 等EMNLP 2023 · 被引用 129 次
- s1: Simple test-time scalingNiklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li 等EMNLP 2025 · 被引用 33 次
- Sampling-Based Approximations to Minimum Bayes Risk Decoding for Neural Machine TranslationBryan Eikema, Wilker AzizEMNLP 2022 · 被引用 10 次
- AFRIDOC-MT: Document-level MT Corpus for African LanguagesJesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet 等EMNLP 2025
- ReMedy: Learning Machine Translation Evaluation from Human Preferences with Reward ModelingShaomu Tan, Christof MonzEMNLP 2025
相关 Paper
- Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation RefinementYichen Dong, Xinglin Lyu, Junhui Li, Daimeng Wei 等ACL 2025
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM AbilitiesDavid Stap, Eva Hasler, Bill Byrne, Christof Monz 等ACL 2024
- Improving Iterative Text Revision by Learning Where to Edit from Other Revision TasksZae Myung Kim, Wanyu Du, Vipul Raheja, Dhruv Kumar 等EMNLP 2022 · 被引用 8 次
- LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question AnsweringRan Zhang, Wei Zhao, Lieve Macken, Steffen EgerEMNLP 2025 · 被引用 2 次
