MENLO: From Preferences to Proficiency - Evaluating and Modeling Native-like Quality Across 47 Languages
Chenxi Whitehouse, Sebastian Ruder, Tony Lin, Oksana Kurylo, Haruka Takagi, Janice Lam, Nicolò Busetto, Denise Diaz
Abstract
Ensuring native-like quality of large language model (LLM) responses across many languages is challenging. To address this, we introduce MENLO, a framework that operationalizes the evaluation of native-like response quality based on audience design-inspired mechanisms. Using MENLO, we create a dataset of 6,423 human-annotated prompt–response preference pairs covering four quality dimensions with high inter-annotator agreement in 47 language varieties. Our evaluation reveals that zero-shot LLM judges benefit significantly from pairwise evaluation and our structured annotation rubrics, yet they still underperform human annotators on our dataset. We demonstrate substantial improvements through fine-tuning with reinforcement learning, reward shaping, and multi-task learning approaches. Additionally, we show that RL-trained judges can serve as generative reward models to enhance LLMs' multilingual proficiency, though discrepancies with human judgment remain. Our findings suggest promising directions for scalable multilingual evaluation and preference alignment. We release our dataset and evaluation framework to support further research in multilingual LLM evaluation (https://huggingface.co/datasets/facebook/menlo).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f68c6a29-9251-4592-85bb-8b297bfff371Builds on12
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng et al.EMNLP 2023 · 91 citations
- HYDRA: Model Factorization Framework for Black-Box LLM PersonalizationYuchen Zhuang, Haotian Sun, Yue Yu, Rushi Qiang et al.NeurIPS 2024 · 79 citations
- Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal et al.ACL 2023 · 41 citations
- Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual AlignmentZhaofeng Wu, Ananth Balashankar, Yoon Kim, Jacob Eisenstein et al.EMNLP 2024 · 29 citations
Related papers
- M-RewardBench: Evaluating Reward Models in Multilingual SettingsSrishti Gureja, Lester James Validad Miranda, Shayekh Bin Islam, Rishabh Maheshwary et al.ACL 2025
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- mR3: Multilingual Rubric-Agnostic Reward Reasoning ModelsDavid Anugraha, Shou-Yi Hung, Zilu Tang, En-Shiun Annie Lee et al.ICLR 2026 · 9 citations
- UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation ParadigmsPeng Lai, Yichao Du, Junchao Wu, Weibo Gao et al.ICML 2026
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward ModelsChenglong Wang, Yifu Huo, Yang Gan, Yongyu Mu et al.AAAI 2026 · 1 citation
