PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, Manohar Swaminathan, Sunayana Sitaram
Abstract
Evaluation of multilingual Large Language Models (LLMs) is challenging due to a variety of factors -the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and the lack of local, cultural nuances in translated benchmarks. In this work, we study human and LLM-based evaluation in a multilingual, multi-cultural setting. We evaluate 30 models across 10 Indic languages by conducting 90K human evaluations and 30K LLMbased evaluations and find that models such as GPT-4o and Llama-3 70B consistently perform best for most Indic languages. We build leaderboards for two evaluation settings -pairwise comparison and direct assessment and analyse the agreement between humans and LLMs. We find that humans and LLMs agree fairly well in the pairwise setting but the agreement drops for direct assessment evaluation especially for languages such as Bengali and Odia. We also check for various biases in human and LLMbased evaluation and find evidence of self-bias in the GPT-based evaluator. Our work presents a significant step towards scaling up multilingual evaluation of LLMs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef724fe9-cb4c-4edc-929d-9059c46e9583Cited by top-tier papers8
- BLEUBERI: BLEU is a surprisingly effective reward for instruction followingYapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh et al.NeurIPS 2025 · 21 citations
- FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and StereotypesJanki Atul Nawale, Mohammed Safi Ur Rahman Khan, Janani D, Mansi Gupta et al.ACL 2025 · 5 citations
- MENLO: From Preferences to Proficiency - Evaluating and Modeling Native-like Quality Across 47 LanguagesChenxi Whitehouse, Sebastian Ruder, Tony Lin, Oksana Kurylo et al.ICLR 2026 · 3 citations
- UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic LanguagesPranjal A. Chitale, Varun Gumma, Sanchit Ahuja, Prashant Kodali et al.ACL 2026 · 1 citation
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot SettingsHamna, Gayatri Bhat, Sourabrata Mukherjee, Faisal M. Lalani et al.CHI 2026 · 1 citation
Builds on16
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu et al.EMNLP 2020 · 232 citations
- Proving Test Set Contamination in Black-Box Language ModelsYonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak et al.ICLR 2024 · 220 citations
Related papers
- IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic LanguagesHarman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari et al.ACL 2024
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMsAli Faraz, Akash, Shaharukh Khan, Raja Kolla et al.ICLR 2026 · 9 citations
- McEval: Massively Multilingual Code EvaluationLinzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin et al.ICLR 2025 · 1 citation
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare QueriesYiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu et al.WWW 2024 · 126 citations
- Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLUFajri Koto, Nurul Aisyah, Haonan Li, Timothy BaldwinEMNLP 2023 · 11 citations
