Test-time Adaptation for Machine Translation Evaluation by Uncertainty Minimization
Runzhe Zhan, Xuebo Liu, Derek F. Wong, Cuilian Zhang, Lidia S. Chao, Min Zhang
Abstract
The neural metrics recently received considerable attention from the research community in the automatic evaluation of machine translation. Unlike text-based metrics that have interpretable and consistent evaluation mechanisms for various data sources, the reliability of neural metrics in assessing out-of-distribution data remains a concern due to the disparity between training data and real-world data. This paper aims to address the inference bias of neural metrics through uncertainty minimization during test time, without requiring additional data. Our proposed method comprises three steps: uncertainty estimation, test-time adaptation, and inference. Specifically, the model employs the prediction uncertainty of the current data as a signal to update a small fraction of parameters during test time and subsequently refine the prediction through optimization. To validate our approach, we apply the proposed method to three representative models and conduct experiments on the WMT21 benchmarks. The results obtained from both in-domain and out-of-distribution evaluations consistently demonstrate improvements in correlation performance across different models. Furthermore, we provide evidence that the proposed method effectively reduces model uncertainty. The code is publicly available at https://github.com/NLP2CT/TaU .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Can Large Language Models Capture Dissenting Human Voices?Noah Lee, Na An, James ThorneEMNLP 2023 · 8 citations
- Selective Label Enhancement Learning for Test-Time AdaptationYihao Hu, Congyu Qiao, Xin Geng, Ning XuICLR 2025
Builds on9
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Disentangling Uncertainty in Machine Translation EvaluationChrysoula Zerva, Taisiya Glushkova, Ricardo Rei, André F. T. MartinsEMNLP 2022 · 6 citations
- Multi-Hypothesis Machine Translation EvaluationMarina Fomicheva, Lucia Specia, Francisco GuzmánACL 2020 · 13 citations
- AETTA: Label-Free Accuracy Estimation for Test-Time AdaptationTaeckyung Lee, Sorn Chottananurak, Taesik Gong, Sung-Ju LeeCVPR 2024
- Uncertainty-Aware Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng et al.EMNLP 2020 · 20 citations
- UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language ModelsRoman Vashurin, Maiya Goloburda, Preslav Nakov, Maxim PanovEMNLP 2025 · 1 citation
