Lune

ICSE2026顶会

VADA: A Multicultural Benchmark for Value-Aware Data Generation and Alignment Evaluation in LLMs

Zhenlun Zhang, Yang Feng, Shihao Weng, Yining Yin, Jincheng Li, Jia Liu

2026年份

摘要

Large language models (LLMs) now sit inside an expanding set of intelligent software systems, from education assistants to legal advisory tools. When these systems ship to global users, ensuring their behavior respects diverse cultural values has become a pressing challenge for software engineers. Traditional software engineering methods such as unit testing and formal verification work well for functional requirements, but they struggle to encode and validate normative expectations like cultural sensitivity or moral appropriateness. To address this gap, we introduce VADA, a framework for testing and benchmarking LLMs under multiple cultural value systems. VADA builds test cases through a modular pipeline that generates scenarios and questions grounded in culturally defined value dimensions. The resulting suite covers 25 value dimensions drawn from China’s Core Values, EU’s fundamental rights framework, and Islamic ethical frameworks. VADA uses a Bayesian ensemble and assigns each evaluator a dimension specific weight based on its observed reliability to evaluate model responses, and we additionally fine tune a lightweight supervised evaluator using VADA generated data, offering a fast and scalable alternative to running multiple large evaluators. Using this pipeline, we build a large benchmark containing 11,865 automatically annotated cases, along with a human labeled subset of 1,000 instances for validation and evaluation. VADA substantially outperforms current prompt-based evaluators, reaching 93.6% accuracy and showing strong alignment with human annotations. Ablation studies indicate that both evaluator diversity and reliability based weighting contribute to performance. VADA also supports comparative audits of recent LLMs and reveals inconsistent alignment across cultural dimensions. Together, these results suggest that VADA provides a robust basis for value aware evaluation and supports the development of culturally aligned LLM based systems.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖