VADA: A Multicultural Benchmark for Value-Aware Data Generation and Alignment Evaluation in LLMs
Zhenlun Zhang, Yang Feng, Shihao Weng, Yining Yin, Jincheng Li, Jia Liu
Abstract
Large language models (LLMs) now sit inside an expanding set of intelligent software systems, from education assistants to legal advisory tools. When these systems ship to global users, ensuring their behavior respects diverse cultural values has become a pressing challenge for software engineers. Traditional software engineering methods such as unit testing and formal verification work well for functional requirements, but they struggle to encode and validate normative expectations like cultural sensitivity or moral appropriateness. To address this gap, we introduce VADA, a framework for testing and benchmarking LLMs under multiple cultural value systems. VADA builds test cases through a modular pipeline that generates scenarios and questions grounded in culturally defined value dimensions. The resulting suite covers 25 value dimensions drawn from China’s Core Values, EU’s fundamental rights framework, and Islamic ethical frameworks. VADA uses a Bayesian ensemble and assigns each evaluator a dimension specific weight based on its observed reliability to evaluate model responses, and we additionally fine tune a lightweight supervised evaluator using VADA generated data, offering a fast and scalable alternative to running multiple large evaluators. Using this pipeline, we build a large benchmark containing 11,865 automatically annotated cases, along with a human labeled subset of 1,000 instances for validation and evaluation. VADA substantially outperforms current prompt-based evaluators, reaching 93.6% accuracy and showing strong alignment with human annotations. Ablation studies indicate that both evaluator diversity and reliability based weighting contribute to performance. VADA also supports comparative audits of recent LLMs and reveals inconsistent alignment across cultural dimensions. Together, these results suggest that VADA provides a robust basis for value aware evaluation and supports the development of culturally aligned LLM based systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 97d658d2-6013-4d5b-89fe-456d8adf01afRelated papers
- MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu et al.ACL 2026
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value DifferenceJing Yao, Shitong Duan, Xiaoyuan Yi, Dongkuan Xu et al.ICLR 2026 · 4 citations
- Measuring Meta-Cultural Competency: A Spectral Framework for LLM Knowledge StructuresSougata Saha, Madhur Jindal, Saurabh Kumar Pandey, Mahardika Ihsani et al.ICML 2026
- ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language ModelsYuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang et al.ACL 2024
- Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-SortsJingting Zheng, Yuqi Ren, Linhao Yu, Yongqi Leng et al.ACL 2026
