DiscoX: Benchmarking Discourse-Level Translation in Expert Domains
Xiying ZHAO, Zhoufutu Wen, Zhixuan Chen, Jingzhe Ding, Jianpeng Jiao, Shuai Li, Xi Li, Danni Liang, Shengda Long, Qianqian Liu, Xianbo Wu, Hongwan Gao
摘要
The evaluation of discourse-level translation in expert domains remains inadequate, despite its centrality to knowledge dissemination and cross-lingual scholarly communication. While these translations demand discourse-level coherence and strict terminological precision, current evaluation methods predominantly focus on segment-level accuracy and fluency. To address this limitation, we introduce DiscoX, a new benchmark for discourse-level and expert-level Chinese-English translation. It comprises 200 professionally-curated texts from 7 domains, with an average length exceeding 1700 tokens. To evaluate performance on DiscoX, we also develop Metric-S, a reference-free system that provides fine-grained automatic assessments across accuracy, fluency, and appropriateness. Metric-S demonstrates strong consistency with human judgments, significantly outperforming existing metrics. Our experiments reveal a remarkable performance gap: even the most advanced LLMs still trail human experts on these tasks. This finding validates the difficulty of DiscoX and underscores the challenges that remain in achieving professional-grade machine translation. The proposed benchmark and evaluation system provide a robust framework for more rigorous evaluation, facilitating future advancements in LLM-based translation. Our data and code are available at https://github.com/ByteDance-Seed/DiscoX.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsSeungone Kim, Jamin Shin, Yejin Choi, Joel Jang 等ICLR 2024 · 被引用 468 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
- Large Language Models are not Fair EvaluatorsPeiyi Wang, Lei Li, Liang Chen, Zefan Cai 等ACL 2024
相关 Paper
- When Does Translation Require Context? A Data-driven, Multilingual ExplorationPatrick Fernandes, Kayo Yin, Emmy Liu, André F. T. Martins 等ACL 2023 · 被引用 10 次
- Quantifying the Impact of Translation Errors on Multilingual LLM EvaluationKlaudia Thellmann, Bernhard Stadler, Michael Färber, Jens LehmannACL 2026
- DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain TranslationZhibo Man, Yuanmeng Chen, Yujie Zhang, Jinan XuEMNLP 2025
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang 等EMNLP 2023 · 被引用 129 次
- EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error CorrectionJingheng Ye, Shang Qin, Yinghui Li, Xuxin Cheng 等AAAI 2025 · 被引用 3 次
