Extending Automatic Machine Translation Evaluation to Book-Length Documents
Kuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang, Ping-Chun Hsieh, Wen-Chih Peng, Vitaly Lavrukhin, Boris Ginsburg
Abstract
Despite Large Language Models (LLMs) demonstrating superior translation performance and long-context capabilities, evaluation methodologies remain constrained to sentencelevel assessment due to dataset limitations, token number restrictions in metrics, and rigid sentence boundary requirements. We introduce SEGALE, an evaluation scheme that extends existing automatic metrics to long-document translation by treating documents as continuous text and applying sentence segmentation and alignment methods. Our approach enables previously unattainable document-level evaluation, handling translations of arbitrary length generated with document-level prompts while accounting for under-/over-translations and varied sentence boundaries. Experiments show our scheme significantly outperforms existing longform document evaluation schemes, while being comparable to evaluations performed with groundtruth sentence alignments. Additionally, we apply our scheme to book-length texts and newly demonstrate that many open-weight LLMs fail to effectively translate documents at their reported maximum context lengths.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0763f8c-7dea-4acc-a410-46d4ff319284Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang et al.EMNLP 2023 · 129 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World LiteratureKatherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray et al.EMNLP 2022 · 23 citations
- L-Eval: Instituting Standardized Evaluation for Long Context Language ModelsChenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao et al.ACL 2024 · 6 citations
Related papers
- Improving Long-Context Translation via Self-Supervised Dual LearningShanbo Cheng, Shuaijie She, Yu Bao, Jianbing Zhang et al.ACL 2026
- LooGLE: Can Long-Context Language Models Understand Long Contexts?Jiaqi Li, Mengmeng Wang, Zilong Zheng, Muhan ZhangACL 2024 · 32 citations
- DelTA: An Online Document-Level Translation Agent Based on Multi-Level MemoryYutong Wang, Jiali Zeng, Xuebo Liu, Derek F. Wong et al.ICLR 2025
- Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length ContextsYuho Lee, Jiaqi Deng, Nicole Hee-Yeon Kim, Hyangsuk Min et al.EMNLP 2025
- MGAL: A Multilingual Granularity-Aware Long-Context BenchmarkChunhan Li, Chenglin Xu, Zongyang Zhang, Jiale Liu et al.ICML 2026
