Dialect-robust Evaluation of Generated Text
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann
Abstract
Text generation metrics that are not robust to dialect variation make it impossible to tell how well systems perform for many groups of users, and can even penalize systems for producing text in lower-resource dialects. In this paper, we introduce a suite of methods to assess whether metrics are dialect robust. These methods show that state-of-the-art metrics are not dialect robust: they often prioritize dialect similarity over semantics, preferring outputs that are semantically incorrect over outputs that match the semantics of the reference but contain dialect differences. As a step towards dialect-robust metrics for text generation, we propose NANO, which introduces regional and language information to the metric's pretraining. NANO significantly improves dialect robustness while preserving the correlation between automated metrics and human ratings. It also enables a more ambitious approach to evaluation, dialect awareness, in which system outputs are scored by both semantic match to the reference and appropriateness in any specified dialect. 1 https://ewave-atlas.org/parameters/ 62#2/7.0/7.9 2 For simplicity we do not include the reference in this definition. A corpus-level reference-based metric could be defined as 1 N i mi(yi) with mi(yi) = δ(yi, ri), with ri indicating the reference for example i and δ : Y × Y → R.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Multi-VALUE: A Framework for Cross-Dialectal English NLPCaleb Ziems, William Barr Held, Jingfeng Yang, Jwala Dhamala et al.ACL 2023 · 15 citations
- On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar et al.ACL 2023 · 10 citations
- Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting ModelChantal Amrhein, Florian Schottmann, Rico Sennrich, Samuel LäubliACL 2023 · 7 citations
- DADA: Dialect Adaptation via Dynamic Aggregation of Linguistic RulesYanchen Liu, William Barr Held, Diyi YangEMNLP 2023 · 6 citations
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningShivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson et al.ACL 2024
Builds on5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- VALUE: Understanding Dialect Disparity in NLUCaleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson et al.ACL 2022 · 57 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Automatic Machine Translation Evaluation in Many Languages via Zero-Shot ParaphrasingBrian Thompson, Matt PostEMNLP 2020 · 7 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English DialectsJason S. Lucas, Matt Murtagh-White, Ali Al-Lawati, Uchendu Uchendu et al.ACL 2026 · 1 citation
- RoMe: A Robust Metric for Evaluating Natural Language GenerationMd Rashad Al Hasan Rony, Liubov Kovriguina, Debanjan Chaudhuri, Ricardo Usbeck et al.ACL 2022 · 15 citations
- Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive SummarizationItai Mondshine, Tzuf Paz-Argaman, Reut TsarfatyACL 2025 · 6 citations
- Language Model Augmented Relevance ScoreRuibo Liu, Jason Wei, Soroush VosoughiACL 2021
- DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to ModelsNiyati Bafna, Emily Chang, Nathaniel Romney Robinson, David R. Mortensen et al.ACL 2025
